Transcription
In the topic of AI, in the section on learning symptoms for use and development, yes, and then there is the adjustment to living together. In this era, I have six topics in total. There will be the subject of meaning.
And then in the part of the types of AI, we separate them according to their abilities and evolution. Yes, and the most important thing is keywords. If we are interested in this field of study, we will find the words. What are some of the things that are actually involved? There are a lot. The teacher has chosen the ones that he often encounters. Let's get to know it and how to use it.
In many industries, what we use AI for, in the end, will be about ethics. And then there is the ethics in the creation part, and then usage. Let's look at the first one, meaning. Actually, there is a definition of meaning. The meanings are very diverse, teacher. Choose the one that seems to be relatively clear.
It's clear in the part about artificial intelligence. It will be something that we create, no matter what. It will be a machine, a computer, or something, or a robot that can work. Imitate the workings of the human brain. What can I do? Whether it's studying, knowing the data analysis, generating new data, and now doing the part about understanding the language of decision-making.
And also being able to multitask. It's automatic, all in all, for that. It will help us, help people to be able to do some things much faster. Difficult tasks get better in the calculation part. For example, in the category section, if we divide it according to our current abilities, it's the last part of ANI. ANI means artificial narrow intelligence.
It's an AI that performs specific tasks. We created it to detect people in images or functions as a translation. Like this, we will call it ANI, which means being good at something specific. But if it's a bit more general, you can do quite a few things equal to humans, we will call it AGI. And eventually, it will be superhuman skill.
At present, it is not yet this era. In the third era, it has not yet reached. I am an ASI. The evolution will start from the body. This is an artificial neural network, which is a model. Mathematics that simulates human functioning. Yes, it is a thesis or doctoral dissertation. Yes, it was introduced in 1943 and has continued to evolve until now.
We've come to the academic conference in 1955. They coined the term artificial intelligence. Let's start. The first time I saw the term artificial intelligence being used was in 1955. After that, it started to be used. Developed the first chatbot in 1965, it was called Izan. Talk to your computer via text. Let's talk by typing on the keyboard.
Language processing is one of the first things. It's the era of industrial application. And then there are the first virtual assistants. The ones that come in are the ones we know. Siri is a voice assistant that was introduced in 2011. Smartphones are now used for many things.
So whether it's the recognition system or the driving system, automatically move the system language translation, transportation, financial systems, these things will be correct. It is embedded in all keyword circles. There are many words if we study the AI aspect. I've found many, so I'll just choose just eight. You'll probably come across them often.
The first word is ML, Machine Learning. Let me explain what it is. Next, if we break it down a bit, we'll call it DL, Deep Learning, and the inside is... As for its internal working character, we will use ANN, which was born in 1943. This is the first one that was born. Yes, and it is currently used in the matter of understanding human language, called NLP.
And in terms of vision, no matter what, it's a computer machine or robot. It's a CV job, Computer Vision. And the one we use, let's create a new message in the composition. In this article on image generation, we don't just call it AI. We call it Gen AI. Gen AI doesn't mean the era of AI. It means AI that can generate something.
It will generate text, generate video, generate images, which will be done later. After that, there will be an example of generating an image for you to see. And in the case where it can do more than one thing, we will call it this one, last multi-model. We will call it LMM, which is this one in current.
What comes out will be money, which will come out soon. The teacher will also tell you a little bit about the usage of Gen AI. And another thing that we need to know is that the word that is often used is the word "prompt," which is a word that we write to command AI to work as we want.
The first one is ML. ML is a subset of AI. If AI is the top, then it will be ML. ML is Machine Learning. Yes, it is a collection of various techniques to create an AI system. So we're going to give you the computer. Yes, in learning various information and processing and analyzing the results according to the information we teach.
Wait, this one, and in ML itself, there are many methods, one of which is something that has boomed since around 2000. It is Deep Learning. Deep Learning is the use of the neural network. It is a mathematical model that simulates the workings of the human brain. It's a filling that's strung together in many layers.
When it's arranged in 10 layers, 20 layers, 100 layers, 200 layers, it's arranged like this. It will be a deep sorting, so it is called Deep Learning. Therefore, the origin of the term DL comes from the structure. The interior of the model has many layers of components. This is how deep it is.
In fact, three layers can be called Deep Learning. Three layers are considered deep. But actually, before this word is called, it will be around 100 floors or more before it is called. This word, but at present, four floors, five floors, this word can be used as a whole.
And the one that inside Deep Learning is ANN, which is the model. The digestive tract that makes decisions. What we want to analyze is what is it in the case of this type of work? It's classified, so if everyone ranks it, AI is on top. Let's try to go down to ML. The size inside ML already has DL. The inside of DL has ANN inside.
Therefore, the sub-unit that we use is ANN, which is the small unit used. Quite a lot of them are still in use today. Even though it's been around since 1943, it looks like this, right? The fake historical news inside will be a mathematical model that simulates that if there is an input, it will pass through a multiplier in this case.
It's called a weighting factor. It's a factor that determines the importance of each input that comes in. That's useful for sorting. What about that? Suppose we want to separate two types of fruits, what should we use? For example, using length and weight like this shows that the input that runs in is the length.
The weight of the inner node will act as a permit that if there is this much weight, enter. If it's this long, it should come out through a channel. Which one is on the Apple Channel or Papaya Channel? Well, if we separate the two types, then it will be like a door that opens and closes to determine the flow of input to the output according to the type we want.
If we enter a small weight, the length a little bit might be a small fruit. If we enter a lot of weight, then we enter the length. If there are too many, you might say something like papaya. I didn't say Apple. The concept is something like that.
Next, let's talk about NLP. NLP is used to process human natural language. Yes, no matter what the printing language is, right? Is this a message or a spoken language? The spoken language or the language we use in this world, there are about 100 languages used all over the world. Therefore, NLP is one of the first uses of AI.
As for CV, it boomed after NLP because NLP will process the text and turn it into printed text. Enter the text converted from speech. There will be details in the processing, right? There are a lot, but if it's computer vision, what we feed in will be images.
The size of the image also has an effect. Use a small camera, large camera, high-resolution camera, video camera. What kind of camera is this? It will make a noise. And it started to be used more in this era. Powerful computers and fast internet, these things.
This will start to boom in the Computer Vision section. As for Gen AI, what can be called Gen AI? If we use an application or we use some program that can create information, the ability to create images creates videos. We can call it generative AI, but currently, it is a catchy word, right? Gen AI, something like that.
Most Gen AI is developed from this multi-model, in the part of generating text, voice, and images, which is what they call. The model is large because it uses a data set or sound database, image database. There is a huge amount of text data to train this AI, so the camp, as he already has a solid foundation in terms of data usage, whether it's from Facebook, Microsoft, or Google themselves, they will be able to develop a better last model because they have a lot of data.
There are a lot, but if we are new researchers, and we do it ourselves, that's the hardest part of it. The first thing to start with is to collect data to have enough to teach the AI. That might be a lot of work. The next thing is the input data.
Finally, the word "prompt" is what we will get. I heard this about 10 years ago. It's quite loud, right? We'll use it to specify what we want from AI. There will be two components: keywords. We can type keywords that we want this to be in chunks. Keyword block 1, block 2, block 3, or we can arrange them into long sentences, such as: "Please draft this letter and that for me."
Or "Do some decorating for me." You can do it like this or you can feed it in chunks. You can dress up like this before Children's Day. Yes, the NLP, which is the LMM model, will process it itself to see what we expect from the user. What do you want?
If we add more details, there are two types. We will get an output that we are satisfied with. It has complete details. Another thing is that AI will be confused if those inputs are contradictory. That's really true. And we want the output to look like this.
In any case, there will be a science called prompt engineering, which is an engineer who is responsible for giving commands to AI to generate the best possible output. The result we want is learning. It's a serious matter that we will write.
How is it that many AIs that are released these days support the Thai language? Now you can type the command in Thai. Let's take a look at an example of how to use it. In what areas do we currently use it? Yes, let's take something close to home.
The first level will be: in terms of identification, such as Face ID systems, etc. It is used for unlocking, lock access verification, identity verification to use various mobile devices. Yes, it is the one that we use almost every day.
Right, in terms of social media, there will be an AI system embedded in the feed. Advertisement news, sometimes we want to eat something. When I rested, it bounced back up. It went to eavesdrop. I'm not sure about sending emails.
Whether it's drafting email messages or filtering or classifying this as a normal email or if it's spam, there will be a screening checkpoint. The first filter is done for us in the search engine section. We search for words. Here are the keywords in the section.
The results shown are close to what we search for. It's the first part that shows up. And we can't find it or we can search for the word. New and more detailed information is also available in the section of the person. Help me, many people may have used it.
Those who are Apple users will use Siri. If it is Android, it is popular to use Google Assistant. If it is Amazon, it will be Alexa. This one will receive commands. Our speech is converted into text. Then follow that command for us in the NLP section, whether it's language translation work.
Gamma correction of the proposed sentence that is truncated. Please give us a new section on analyzing the sentiment in the text. The fact that the author is writing this now shows that he is dissatisfied with the food menu we commented on. If it is a form of insulting on social media, these people will be screened for their part.
Sentiment analysis, or it is a chatbot that responds when buying or selling things. It has AI embedded in it. This is part of the Computer Vision work, such as detecting objects in moving images. Is it moving or still? Object detection, we must want it, no matter what it is.
Like car detection, smoke detection, these lights, and in the matter of analyzing the image to see which products this batch of products can be used and should not be rejected. Or should it be sent for sale? Analysis of security in the CCTV section. If you turn these off, you will need to use a lot of visuals.
And what is currently popular is for image creation. It is image generation. Print the printed image. We type the message we want now. He either gives us the picture or we add the picture. There are both cars and people, so we typed to help delete the cars.
Give it to me like this and it will deduct the discount for us. And then only the people in the picture are left. Nowadays, it is much easier to use. Most Photoshop has some features that they have built AI into. The travel part is calculating the route.
The shortest one for us will be the various maps, which will have AI to help manage. This is about checking for unusual transactions. Yes, the withdrawal of the abnormal transfer in it's a matter of money and time. And then in terms of frequency, it will be good.
In terms of medicine, it will help a lot. For the analysis of photographs, such as x-rays and MRIs, etc., they will help us pinpoint the target. These are much easier to analyze. In terms of business itself, it starts from business planning.
We can use Gen AI to design the structure of the business plan. The marketing plan is naming. These product design shops can do it. In the production line, it starts with cost calculation. Yes, to inspect or QC the product before it is for sale in response to market demand.
In setting sales targets, these can be done in agriculture. It may be used in analyzing various data from sensors. Various ways to care for the health of crops. The ones we plant have both the use of images in the expected amount of harvested rice or look at the current condition of the plants.
Is this normal or is it a disease? Yes, I will go and help with legal matters. Yes, whether it is searching for a court case examples or words similar to what we are interested in. This might help. The story of data analysis simulation science or create a mathematical model that helps in data analysis can also help a lot.
And then in terms of planning learning and doing activities for children or students or in developing teaching media, we can print. You can go in and help design activities for 2-year-old children. The activity duration is 30 minutes. They will type and report back to us what activities we should have.
What did you play during those 2 hours? Yes, we can go shopping and choose from there. As for the research part, if we go to study, master's and PhD degrees are very helpful. Right, we throw out papers or research that are when we enter this article.
We type that summarize the article in five lines like this. He will summarize the main points for us. Or it could help in rephrasing words. In the part of reading the analysis of research that we are missing something that needs to be typed in. You can see it directly in the development section.
Software, the most famous one is Copilot, which helps from the beginning. The beginning of writing program code in the part about throwing the code in for them to do. I wonder if there is a problem with the code or not. In the section on throwing working codes, let him rewrite it in a new format.
That can be shorter and more concise than before or we have a sample code in Python language and ask him to change it to C language. It can be done. It can convert across languages in the movie and production section. These games are similar in terms of visual effects, character design, and script design.
Movies, these can be imagined. We have run out of ideas. We might ask for three ideas and they will present them to us. We will like or dislike one thing and then after shopping, let's look at the characteristics.
In terms of usage, if we define it in the form of input, right now, the input that we give to AI will be text and images. Audio, video, and various sensors, such as IoT work, will receive these five things and process them for us. As for the output, we can choose whether we want the output to be text or an image.
Most audio and video outputs nowadays come out in these four ways. Choose what you want to see in the first part. Before the teacher arranges it like this, the first part will be the voice input type, the text input type, and then the image input type.
Let's look at the voice input type first. What's interesting? Let's use Voice Assistant sometimes. We're familiar with calling it VA. If we search on Google, we can search for VA and it will be there. There are many brands and services for us to choose from.
Most of us use it through smartphones, no matter what. It is a call, flashlight, and command. Searching for music, ordering YouTube, and all that stuff can be used right away, right? What can I use? Google Assistant, as the name suggests, is Google.
This is definitely from Google. Can be installed on both Android and iOS. Can be installed for use in ordering. Hardware in the data request section. Various weather conditions, you can do it or not. It is also an order to open another app.
Another camp is used quite a lot. Yes, Irina. Irina turned on the flashlight like this. I use it often. Has anyone used it for a while? It won't be like that at first, it will be like listening to a sound. Many people have heard the teacher say that Prime Minister Kai said that too.
It works like a buffalo, a snake works, but we use it for a while. One thing is, it's going to be like this, only one person can say it. It's like it's starting to remember something. Something that is an accent of the people who use it often, it's other people who come in and use it. Our device will not respond very much.
Next, Alexa, oh, the Amazon is designed to not be integrated into smartphones. I'll make it into a box and put it in my house. Then tell Alexa to open the file. This and that, turn on the air conditioner, it will be like this. It is a type of box that is placed at various points in the home, put it in the living room, kitchen, etc.
This is helpful. If you have any questions, just ask. Yes, this is from Amazon. Next is Cortana. Many people don't use this one very often. The teacher said that most people will not turn on this mode when they install Windows. Windows 10, Windows 11, many people will turn off this mode or some it's been closed since the beginning.
We can actually use the control in Windows. It will be embedded in this. Try it. Go home and open it. Personally, Bixby is from Samsung. Anyone who uses this brand will use this to receive commands. The sounds may be similar, but the design is different.
And the camp and the language accent that they answered. If we talk about it, we'll be from different camps. Like this, whichever brand you like, use that brand. Okay, another one that many people might like to use. Yes, it's a way to search for songs by voice.
This will be an analysis of the sound that we really want to know what song is playing in. What song is this wave? We open the app. When it comes up, you can actually search on Google. You can search right away. There will be a Google microphone icon.
Then we let it listen to the sound. After about five seconds, it will pop up and say: What song is it? Whose is it? But if it is the application itself, it is popular as Shazam. This one will receive the sound. Music that comes from speakers, from TVs, radio waves take a while to search.
It will soon come up with who this song belongs to, who is the singer, which album, and year is it from? But if we want to, we can't think of it, but we can do it. Do you have the melody or can you sing it as lyrics? That's right, I'll use this Saho.
We can make this into a melody that goes like this. Well, it's still possible to search for what we need. What was that song just now? This one has a bit more gimmicks than the original. In the past, it was about reading or singing.
You can go in and that is voice input. Now, let's look at text input. There are two major companies: OpenAI and Google AI. This one will be used as a large language model, last language model. Let me first explain the chat GPT.
The chat GPT, if it is the free version that we use on our website, is model 3.5, but if anyone loses it, the money that he gives will be used as a 4.0 model. It will be from the company OpenAI, which is open source. What can chat GPT do? The first thing is to answer.
Chat with us, type a message and chat with GPT. Sure, we can chat with friends instead of talking to each other. We can chat with GPT or create various content such as scripts and articles. Requesting program code or song lyrics or write a song with a slogan for Children's Day or something.
You can do everything in the question-answering section. Yes, we tried to answer, but we were satisfied. With that answer, we also choose another thing. Yes, most of it is writing code. The one that is really good is the Python language.
Let's try to go down to Java. There are two of these that are quite popular for requesting Python code to calculate this or that. You can do something like this or we can tell you that we will start building a detection model with pictures like this.
We don't know where to start. Anyway, we type into the GPT chat and ask for the code to create an AI model to detect and capture images of dogs and cats. It will then give us a code along with a link that is a reference in the section.
The translation of the teacher's language is that it is currently quite the sides are about the same and I don't really care about the model. Which model is more prominent than the other depends on with us, we are satisfied with the language you have translated. We are another layer of filtering.
Next is Google Gin. I secretly support this one. Right, because I just came here and it's not true. Just arrived, got a new look and a new name. New from the original name, launched last December, I was going to. It is a larger model than GPT chat and has more capabilities in terms of receiving input.
If GPT chat, we can only type in the text. But if it's Gin, we'll send you a picture. You can also have him analyze the image for us. Is anyone there or are we making noise? Go in and let him analyze what it might be like.
What sound is noise or what is sound? What kind of environment is this? So what will it be like? A model that is larger than GPT chat. What can it do? Basic capabilities, this is similar, right? Whether it's writing, translation, text, compose a song, write code.
Yes, but the latter is not like GPT chat and will have the ability to describe images. And then there's image analysis with this example. Write a story with an elephant and a monkey, one at a time. Long ago, there was a big, kind elephant named Pai.
This guy is the last one to be my friend. Are you in love in the forest or something? You can write the code here. Write the Python code to calculate the area of a triangle here. He clearly explained that this is about receiving input values for calculating the area of a triangle.
And then the processing has comments. With it, ready to run, simply put, run it, guys. This is an example of how to use it. We can use it in two ways. If we use it like this, it's called using the website, but if we do it in the manner of a developer, it's like this.
For our developers, we will use Gin in the back house, which is called using through API. We have to install Python properly, download the generative AI library, and then we log in to request API Key. Google Gin, you can choose to use two models.
The first model is Gin Pro, which receives input as text. The other one is Gemini Vision, which receives input as images. Let's say the teacher asks a question. Well, how many languages can we understand? Then we run this Python code and it will respond.
Let's say that we can understand 100 languages. It will report to us all today. This will be the way we use the code. You can take this code and build it through an app, ILD, or a web application. Because we are doing it at the ground level, we can't be used through a website or by drafting a document.
I'm taking a vacation to Central Nakhon Pathom on April 2nd. I'm taking a vacation today, right? He will draft a letter or any drafting. Let's come in the form of a tent. Here, put it in. Our name, our address, email address.
If you are already there, it's the name of the manager or the person we want to leave, and then he will draft it as follows: "Regarding my respectful leave, what is my name and position? I would like to request permission to take leave to go to Central Nakhon Thom on April 2nd. I will return to work as usual on April 3rd. I have prepared necessary work. It's all settled down, let's get started."
You can also add anything you want. Another last respect. It's a language model as well. Wit.ai is suitable for NLP work. It can input a lot of text. Process the text we type for us. What does going there mean? What did we ask for?
This one is from Facebook. We can use our Facebook account to log in to Wit.ai and perform the following functions: By understanding the language we speak, we may be able to develop voice-activated IoT devices. Like this, turn on the fan at number 2 like this. Wit.ai will separate it for us, the user.
What does the job require? Open it, open anything, it will be done. The pile is the word fan followed by an adjective, which is number 2, so it will go and order. The hardware is provided to us in the visual aspect. Yes, we will call this work in computer vision.
Computer vision is a subset of AI. It is a machine that receives input in the form of images or videos. See the understanding and interpretation of what is there. There are many works inside those pictures. The teacher has selected eight main works for us to see.
You will often encounter this in your first daily life, right? It is to classify the picture into what it is, a dog or something. What kind of dog is a cat? What kind of cat is a cat? Yes, object detection is a hit. Please frame where the dog is in the picture.
Where is the car in the picture? Improve the image. We have a dark CCTV camera. In this case, we may make some improvements to the image. Light it up to reveal more details inside. You can go up in the image segmentation section.
It will be like cutting the tint out here. The trees here are mountains, here are buildings. Like this, if we want to focus on the details. Another image of OCR to detect various texts such as advertising signs. The car registration is a water meter, an electric meter, etc.
This will be used in the form of OCR or currently, what is widely used is super resolution to make our images. This is more detailed. We have, oh, the image quality is not very good. Zoom in. Go in and the picture is broken like this.
Another one is the captioning. It will be given. Computers know what's there. In some pictures, the last one is image generation. Many of you may have used it in many applications that create images as we like.
Let's take a look at the first one, the image classification. Yes, image classification is not broken down into smaller pieces. Go to the inside section, it will show you the whole thing. Yes, a cat is like a cat. Here is the inside of a popular image classification process.
Use it, it will be a CNN model. Inside it will consist of many image filters. I'm going to filter from the original image. Continue until you reach the most important features and you can answer them. So this is a car. This is a motorcycle, a traffic sign like this.
Inside, it will work in the form of linked together step by step like this. Right, we will analyze in sub-sections. Each subset of the image is extracted. This one is important, the other one is not important, just throw it away until you leave.
Let's say this is a picture of a bird, for example. In the object detection section, we will look deep into the inside to see that there is a dog, two cats, and one little duck, something like that. You don't look at the whole picture as an animal, but rather as an inside look.
Let's take a look at what's inside. There are many popular types of chestnuts. Some names may sound familiar, some may not. Yes, there is YOLO, there is Single Shot, there is RCNN, Fast CNN, Faster, Fast Test, there are all of these. Its duty is to tell us that there is an area.
What is in our image is the culprit. Improve this image as a simple example. Well, if we have an image that's too dark, it can be adjusted to the right-hand side. There will be an adjustment. The color tone of the image that we see in the details on the back is this is more clear.
We will call this work "image enhancement." It does something to the image inside, such as adjusting brightness, adjusting contrast, adjusting the color tone and saturation. And the color temperature is warm and cool. Yes, and in the part about emphasizing the edges or making the image from blurry to very sharp.
It can be done if it is a segmentation work like this. We will break it down at the level. It's a pixel. You can see the kittens if you want. These yellow pixels show that the AI knows the region or area. The cat behind is a blue tree.
The sky and over there is the field. Grass like this can be separated into shots. It can be a scene-by-scene type of work. This is something you can do right now on your smartphone. If we turn on the portrait mode, it will be the DOF, which is the depth of field.
The foreground is clear, the background is blurred, it will know. Who is near or far? So the first segment of the camera in a smartphone that it allows us to do is separate two things: the object and the background. There are many ways to do this, such as looking at areas of skin color that are similar in color.
The surface of the object in the image or to see the image. This has the limits of being human. What does a dog look like as a cat? It is the use of edges or some work, some algorithms as well. A combination of both looking at the color of the object and looking at the scope.
It also has physical objects, but at present, it is impossible to avoid them. Let's use AI instead. It's easier. Yes, the traditional method above. Easier efficiency with the core of AI that is the image segmentation task, we will call it DNN.
Let's take a look at this next. A little closer to us if we use it. Enter and exit any place where we have a record of the sign registration or brand recognition of the grille. The front of the car or the driver's face? Those jobs will require many sectors to do one thing.
But if the one that is related to letters, we call it OCR. It's so catchy that you might forget the abbreviation. It's an abbreviation for Optical Character Recognition, which is the recognition of the person. The real characters are not with the light, they are with look at the picture to see what the message is in the picture.
Some of them are English language recognition, Thai or Korean or Chinese. Many people may have used a translation app that sends pictures. Once you enter it, it will convert into the language. Others that we can choose from have OCR embedded in them. They are very popular.
There are three of them that we use. So if we are good at writing Python, we can import three libraries. Let's try it out. There will be a Test Slag, Easy OCR, and Pad OCR. These three are open source and can be used without any damage.
Tang is the one that works quite well. For Thai, we will use Easy OCR. For example, the teacher uses Gin. The teacher threw this image in. Pattani, number 43, threw it in. Gemini answered. That Pattani Pattani English Thai language.
He found it, he found it here again, ah18, and then he found this one, as2, but he didn't find the number 43 or the number 4. Today, Gemini can throw in pictures. So it will come out and say what is this? Many people are doing research on this.
Because we are still developing and improving continuously. Suppose we have a picture like this, have we taken it? We have a fairly small image source file, like 100x100 pixels. It can upgrade the resolution to a level. This is what this job is called.
Increase the resolution of the image to super resolution. Some people will use it in the first stage. Yes, let's say that the images from the CCTV camera have poor quality when they are taken. What is it? Let's assume that you take a picture of a license plate from a distance.
When it comes, it will be a broken image like this. We will send it to OCR to recognize that it is a registration number, 7788. It might give the wrong answer. So the first level that he does, he will do this level. Before he came and upgraded the resolution, gradually throw this clear image into the OCR, and it will be able to recognize the characters.
It's much cheaper. It has to pass the super resolution level first. The algorithm is the starting point. They will use a cubic sheet. The adjacent color values are processed. And then average it out and stuff it in the case that we zoom in to make the image larger.
But now it will be this one. In the AI circle, we call it a generative adversarial network. It is responsible for generating image details. Come up to make it more realistic. It will be there is a slight disadvantage to adding it if we have a picture of the leaf.
People's faces are not very clear, like those with a lot of numbers. Maybe the original didn't know what the nose looked like. How is your mouth? How are your eyes? What will happen to your mouth? There's a little bit of bias that will add some embellishment to it.
It might be an unrealistically beautiful mouth or a beautiful nose. Exaggerated to make us look better than the original. Something like this might make it possible for us to do that. So the limitation of its use is if the input is not enough to leave any traces.
Yes, it will be enhanced by AI. It will be a bit exaggerated. Sometimes our face is too white and our nose is too big. What kind of exaggeration is this? It could happen. Using these will be like this in the captioning section.
This one, Google Gin, does it well. This is an image. As for the research that they are still researching. Stay, we threw the picture in. The AI replied like this. This is a man in a black T-shirt. I'm playing the guitar like this.
I know who is in the picture wearing the colored shirt. What are you doing? Or what are you doing? He's a construction worker in an orange safety suit, like this, working on the road. Yes, or there will be something like this. There are two women playing with Lego toys.
This type of work is called image captioning, which Gin can do. We can use the code example. Right, in terms of using Gin Pro, it is still free to use as a PR Vision provision, meaning that we can throw in image inputs, teacher.
Throw in one picture and the teacher will type and tell you like this. Please describe this picture for me. Look at the body. Like in the lecture, the teacher threw a picture of a horse. Go in and type in the words that describe the image.
This is the answer that Gemini gave to the area. Yes, a brown horse is running in the fields, so Gemini knows that what is in the picture and what activity is he doing? In the picture, everyone looks at the details. Gemini asked Mana to describe this picture.
Teacher Pim just went in briefly and the teacher threw a picture. Go to Gemini and type back that this picture is a photo of a dog and cat sitting on a carpet. The white animal, the dog, is wearing a brown leather jacket. Looking seriously at the camera while the cat is sitting.
Standing next to him, his eyes closed and looking very relaxed. Well, we probably won't take it this seriously, right? This is used in a descriptive manner if describing. This is the picture when the teacher tried the picture.
It describes a man and two women holding a tablet and wearing pink ties or something. But if we don't ask like that, we won't get it. Caption: We asked like this, everything in the picture is there. How many people?
Gemini, please answer me. Three people. Because I know that there are people inside and how many there are. The next person who can answer this question should look at this type of work. Image Generation, many people will create illustrations for making slides here and there.
We can't imagine that, so please type it. Example of printing, the teacher typed that it was a field. For this flower, we will give you the drunk flower field. There are two things to create and return.
Search for someone who matches, but most of them don't. I won't search and then paste it for us. It will be in the nature of the new creation because the image itself is generated by AI. It will not be copyrighted because it will be considered not duplicating the original, combined into a new image.
Let's call it a group of data. It's actually a group of data. So when we whatever you want, it will give you something that is realistic or not. It would be exaggerated. It doesn't care about us at all.
Example, for example, the teacher just uses this keyword. Eve has a lot of banana trees. Eve also has banana trees. Then come and paste it here for us to see how to make it complete. Ah, as we want, something like this.
Image Generation. Currently, there are many. The camp, right? Whether it's DALL-E, this one is from OpenAI, MidJourney, this one is used by many people, but if we want to create in the mode that is full of his options, it would be a magnificent creation.
It costs money or it will be Disco diffusion and in the part of ImGen, ImGen is from Google AI. It has quite a few other abilities. The teacher will review ImGen for you to see and the picture.
These two images were created using Stability AI and Stability DALL-E. We or it will be Deep Dream Generator, this is also from Google or from Art Breeder. There are many brands to try. We search Google, click on the first link to try.
You can use it to play right away. It's an internal technique in 2015, right? Grant came first. The technique of creating images from the words we type. I went there first, but now I'm going to start.
I don't use it often and will use it as a division model. This review is better because it gives us you can go in and use it and create 100-125 images for free. There will be an installation of the Stability AI Library, which will connect to the API for throwing messages.
Then we receive the image, which Stability AI does not support Thai language, so the teacher has to use another library, which is Google Translate. We type the words "children are reading." The book and the picture are realistic. They look like they are real.
First, change this text to English with Google Translate, then we'll throw it. This message has been changed to English. Then go to Stability AI to create an image of a child sitting and reading a book on the beach.
And then it's kind of like, uh, it's kind of realistic, whatever. Yes, you can run the code. Here is an example. Yes, the man is playing the guitar under the apple tree. Type it into Google Translate and translate it.
The first level, the guitar player is under the apple tree, will take the translated English and throw it into the API. Then he reports back as an image for us, so we will get an image as a square site. Trying to create a guitar player under the apple tree, this will be for us.
Elephants are playing in the waterfall, so we can get some pictures. What we want comes from the conversion with Thai language. Yes, it is in English first and then it will create an image. When we run it, each time we get the same output, we run it the first time and get a picture of an elephant playing.
At this waterfall, we ran around again and we got it. Another action that he made for us. It won't be the exact same as the original. This one is a panda fighting a monkey, it's kind of fantasy. Right, why is the monkey starting to look like he's fighting?
It looks like a panda if everyone looks closely. It's like they're mixed together a little bit confusingly. It should be like this. It's a monkey face and it fights with a panda, but it turns out that the body is considered a cross-species hybrid.
If everyone looks closely, it's a monkey with a panda face. This is a field of banana flowers. For the fancy style, the easiest way is to just play around. Let's just give it to us. It just pops up.
Like this, there are flowers, a cat, and a banana. Come in front and let me know that you already have one. I have everything. This is all that I did. The first thing you can do is generate an image and read the text.
What we typed is the image that was generated for us. It will be a high-resolution image. You can download it and use it without any problems. Copyright, we can have any kind of image we want. Or maybe many people might try this.
You can play it. We've already thrown our image in. We type what we want him to do with it. Our face might be like a horse riding or something. No matter what direction you take it, they will create it.
The character in the picture is a face. We are in a posture or composition. As we typed, it's ready. Yes, this one is called Personal AI image Generator, or multiple people. You can change this right away.
It's a headshot. You can change the hairstyle, but you have to take a close-up shot. And try to make the background not have any complicated backgrounds, take pictures like this. What style do you want for a plain colored wall?
Go ahead and do it if it's good, we'll watch it. No, it's not strange, but if the sync doesn't have a head, we'll watch it. It's a little bit puffy if the components are arranged. Another thing that I can do is not quite right.
This website was recommended by my teacher, so I did it. It can do many things. Text remover. It will detect where the text is in the image. And then it just deletes it for us. Can we delete the text or delete the watermark?
You can or you can delete the background. If you leave it, it will cut out the background or it will be scaled, making the image have larger size, has higher image resolution. We will call it image resolution, but in this work, it will be called upscaling, which is an adjustment.
Clear foreground and blurred background make the front look. It stands out more when compared to the back. It will be a clear front and a blurry back or it will be retouching old photos to make them smoother and clearer.
And then there's another type of work where we have images. Black and white, like the era of your grandparents. When you go in, it will choose a beautiful color for us. Sure, you can adjust the color over there.
It's too dark here, too bright here, you can adjust it. We're taking pictures. The picture is quite blurry. There are two types of blur: blurry and blurry. The number means that we might adjust our focus.
It's not good, there's another number called motion, which is used when we take pictures of something we're in. On the car, we are on the car and we take pictures of the view or we are outside and we take pictures of the car.
Are we filming it moving or are we filming ourselves? It is you who moves that will cause behavior. A little bit of pixel movement, we call it motion, which AI can fix. It will know that we need to move the pixel back.
Where is it that the image has less blur? Or it could be retouching the image. This ImGen can do it in the part of AI that we use in our work that is close to us. It is this one, face processing. The input is our face.
That's it, this is close to you. Ultimately, it's because it's our own face. What can we do with our faces that we use AI models? First of all, it's been a long time coming. Since ancient times, it has been an inspection.
There are many techniques and methods for capturing faces. Oh, sometimes you might only see a straight face. If you tilt your head, you'll meet a stick. If you look down, you'll meet a stick. Yes, they have continued to develop and improve.
Face recognition. Face recognition can tell who this person's face is. Allow the use of that device. Can I enter this room? Another one is age detection. Age detection.
What age range should we be in? It's using various points on the face. Work, whether it's the mouth, ears, nose, or eyes. Now, let's use it in emotion recognition. Yes, in terms of taking care of customer satisfaction.
Why are you eating at our restaurant and walking out? If you look so grumpy, I'll take a look. It can reach satisfaction or it can be called face aging. For example, if we are 50 years old, we can reduce our age to 30, 20 years old or we can increase our age from 50 to 60, 70 years old.
We can call it... This is another one that we use every day: face aging. Right, it will be face swap. I'll show you an example of switching faces. Please show me the face.
We may switch faces with friends, it's okay, just our heads. Friends, let's switch to the first example, which is face detection. The default algorithm only supports straight faces. You don't often see people with faces tilted and looking up.
But now, it's all done with deep learning, and it's all done with AI, whether it's... The face is tilted to the left side. The left and right sides wear glasses, even though they are all seen. Yes, because the trend comes in a rather small face.
There are many different hairstyles that open up the forehead. Open the forehead and you can find the original character, which is HAR CL Kate. Her job is to detect the face from the training. The face will have another reading, which is LBP, to view various features on it.
The face, whether it be the cheek grooves or other grooves. If it happens, we will be able to tell the area. This should be the page, but now it's the one below uses training methods like deep learning to find out like this.
This type of crossbow is not a crossbow. One thing they do in face detection is spoof real faces. It's not a fake face, it's a real face, but it is a real page that does not come from a real person. The real face is from an iPad that was taken from still or moving images.
Here's how to fix it: There are many types. When we scan a face, there will be the voice command came out saying please turn left and right. Bending up and down, which is random, we may have a hard time preparing the video file correctly.
With what it requested at that time, let's say the teacher leaned forward and said, "Please turn left." And then when we look up and do this, it will know. So this is a real person's face, not from video clips or not from printing paper because if you print it, it will probably tilt up and down.
It's not like that. Many people will do that. Do that. This face recognition part will follow face detection, which should be able to tell whose face it is. Yes, who has the right to allow access?
Or this one should be famous. This is an age detection. We can use the AI model that is released as open source to try it out in front of us. Suppose we are 22 years old or something like that and we look at it.
When you go in, it might say you're 30 or something like that, or it might be too young or too old. This is possible if the AI is accurate. The answer is correct. Some AIs don't give a definite answer.
They give a range, such as 20-30, 30-40. It gives a range. Like that, you can choose this one. Like the free version that the teacher came to try. Then, they will detect the faces of these four people and then they will predict that this person is 33 years old, this person is 26 years old, this person is 31 years old, and this person is 26 years old.
Next, let's look at the Landmark page. The Landmark will find a point that will have the version that is not detailed will have about 21 points, but the version that the teacher showed everyone will have 468 points, starting from the forehead down.
To the tip of the nose and eyes, the benefits of the facial landmarks or important points on the face will help in AR work, such as using the glasses. Wearing or applying a mask or adding moles. Whatever it is, it depends on what it does.
Let us fill in the correct posture or gesture on our face. Wherever you are, this is the Landmark, which is called in Thai as the important point on the face. It will come from this information, right? We will put on a mask, put on something like this, it will be an AR that can be added later, which is popularly used, it is called Media.
In addition to the 468 points on the face, what else can Media do? Detect what the object is. This is what Media Thai can do. It is a standard that he can do. Trained to remember the whole picture, what it is, or it could be Hand landmarks.
It will be a point on the hand, the thumb, middle finger, ring finger, and little finger have about 21 points. Or if anyone wants to do it in the hand gesture recognition part is also usable. Media can do this too. Segmentation, for example, selecting people and backgrounds or performing segmentation on specific parts.
Or is it face detection that is aware? The characteristics of the tilt of the eyes, ears, and both cheeks are the landmarks just now. You saw them, 468 points. As for who will do the project? Mini project that knows the whole body, I recommend this one.
Post landmarks, 33 points, eyes, face, shoulders, elbows, knees, and legs go to check for compression. Jumping floor, counting, lifting data or any kind of Thai silk, it will give you a good posture. Of the 33 important points on our body, we can use them to further develop our ability to recognize emotions.
Yes, there will be sadness and fear, but sometimes there will be fear and sadness. It's hard to tell if we're sad or not. I'm scared. I'm happy and indifferent. This is it, we can go and train, we don't have to.
We need to use his database, right? We do it. We might do a sad face like this 300 times to train ourselves that this is the face. It's sad. I'll try to do about 300 pictures. I'm afraid that after doing about 300, we'll train and then we'll do it.
Try it on our own page when we change. The facial gestures according to those emotions are you will be able to accurately remember which emotions your face expresses. Where is it? It's about aging. As far as I can see, there is no free API available for use.
This one is from Disney, you can guess which one. Where is the page that is not filled in the left page? The left side is the normal page. The right side is the added page. It is about 60-70, which is not necessary. You have to move forward only.
You can go backward. Yes, you can go to age 12-14 because that the most important feature is the wrinkles in the part. These various layers are sagging. Muscle tension or fatigue. What will hang first, the cheeks, the eyes, or whatever.
They're going to use those features. Yes, it's called face aging. SW, there will be a technique called "Sing" by the teacher. The possession is there, right? Teacher Sopol is on the side. Right now, the teacher will enter the four people on the left.
It will be a face swap. The right side is taken with a normal webcam. Yes, the concept is that they will take the face, nose, and mouth. 21 important keywords to lock in or fitting into the new face of this person or that person.
This person and this person have a free API. You can try running the code and play. Take a look. It looks like this. Four people transformed into one. Professor Sopol is all gone, this person.
This person, this person, this person is completely possessed. All four bodies have been seized. This is called a face swap. It's almost over. Let's end with the last two topics. If we are going to develop AI, currently the most popular language.
There are about four popular ones. Number one is Python. Number two is Java. Then there is the C language. If it is R language, it will focus on analysis. The information is about research. As for JavaScript, it will be about taking the completed AI and linking it to the web.
This application is about ethics. It has two parts: the creation part and the part that is used in the creation part is this. Any data used to create the model must be authorized. It is clear from the informant that we prohibit it. Secretly stealing other people's data to train your own model.
We in the commercial sector ourselves must be obtaining data with permission. It is written in detail that information can be disclosed. Or you can provide that information in the part of AI that has been used in the part of usage that must be no bias, no bias.
Let's say we go and be our factory's HR department has 1,000 applicants competing for 30 positions in our company. We are too lazy to filter them ourselves and enter them. The features of various properties are filtered by AI. AI must be neutral, no.
It is said that only men are selected. What kind of women are these? They have to work. Be neutral and use non-informative information. And what comes out must coexist with humans can do this without causing division or conflict.
Contrary to humans, it's not like you can use it and give it away. There are various negative effects on privacy when Teacher Kee already talked about getting information. Or is it the data that is produced by AI? Does it violate anyone? This is what it is.
This is something we have to consider when we create. And then when we use it, it's done.