📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

AI Agents Full Course 2025 | AI Agents Tutorial For Beginners | Agentic AI Course | Edureka Live

edureka!7:06:58

Transcription

Hello everyone, and welcome to the AI Agents Full Course. Your step-by-step guide to the future of autonomous AI systems.

In this course, you will start with the agentic AI. Understand how AI agents can reason, plan, and execute tasks independently. We will also cover the foundations of AI and deep learning, understanding large language models, transformers, and NLP. You will also explore LangChain, RAG, MLOps, and prompt engineering while learning how these tools power next-gen intelligent agents.

Finally, we will look at real-world AI advancements like DeepSeek Coder and OpenAI, explore the agentic roadmap, and prepare you with interview questions to kickstart your journey. So, before we begin, please like, share, and subscribe to Edureka's YouTube channel and hit the bell icon to stay updated on the latest content from Edureka.

Also, check out Edureka's Agentic AI certification training. It is carefully crafted to meet industry demands and prepare you for the future of intelligent agents. You will gain practical skills in LangChain, RAG, LLM, MLOps, and more through live instructor-led sessions and hands-on labs. Whether you are a beginner or a tech professional, this course helps you master the concepts and accelerate your AI career. So, check out the course link given in the description box below.

Now, let's get started with our first topic, that is, what is agentic AI. Agentic AI is transforming industries by allowing machines to learn, adapt, and evolve independently. Similar to live organisms, unlike traditional AI, these intelligent agents investigate, optimize, and develop solutions over time without requiring direct human participation. Recent advancements include OpenAI's DeepSeek Coder, which automatically analyzes massive amounts of data to provide detailed reports, and Google's Gemini 2.0, which improves AI's capacity to plan and reason across different data types. ServiceNow's AI Agent Orchestrator is transforming enterprise automation by coordinating many AI agents to address difficult business concerns. As these systems become more powerful, they have the potential to unlock ideas beyond human imagination, ranging from wind turbine blade design to AI-driven company management.

Let's start with our first topic. What is agentic AI? Agentic AI denotes artificial intelligence systems capable of autonomously executing actions to attain designated objectives, unlike reactive AI, which only responds to inputs. Agentic AI is proactive, capable of planning, adapting, and making decisions autonomously. So, let's explore deep into agentic AI and see its capabilities.

Agentic AI is a type of artificial intelligence that exhibits autonomous behavior, enabling it to take actions and operate without continuous human guidance. It is goal-driven, actively working towards achieving specific objectives rather than passively responding to inputs like reactive AI. And with advanced decision-making capabilities, it can evaluate multiple options, select the optimal course of action based on current conditions and acquired knowledge, and adapt its strategies dynamically in response to unforeseen changes in its environment. Moreover, agentic AI demonstrates proactiveness by taking the initiative to act rather than waiting for external triggers, making it highly effective in dynamic and complex scenarios.

Now, let us see its relevance in the current AI market. When AI systems can act autonomously to accomplish predefined objectives, we call that agentic AI, making it highly relevant in the current AI market. Its autonomy allows it to operate without continuous human guidance, making decisions and adapting dynamically to achieve objectives. This capability is complemented by its advanced problem-solving skills, enabling it to evaluate complex situations, strategize, and respond effectively to challenges. However, the growing adoption of agentic AI also raises important ethical considerations, such as ensuring responsible behavior, minimizing unintended consequences, and maintaining transparency in its decision-making processes.

Now that you know about agentic AI, so let us discuss how it differs from other AI systems. Agentic AI differs significantly from other AI systems in its autonomy, decision-making, and adaptability to achieve long-term goals. Unlike reactive AI, which performs predefined tasks only when prompted, such as spam filters or image classifiers, agentic AI takes the initiative and operates independently. It also contrasts with generative AI, which focuses on creating content like ChatGPT generating text, but it is not goal-driven. By combining autonomous behavior, strategic decision-making, and the ability to adapt dynamically, agentic AI stands out as a powerful system designed to achieve specific objectives in evolving environments.

Now, since we know a bit of differences, let us see the comparison between generative AI and agentic AI. Generative AI and agentic AI differ in several key aspects that define their functionality and applications. Generative AI is primarily focused on creation, excelling in output-focused tasks such as generating text, images, or other forms of content. Its adaptability is limited as it relies heavily on prompts for guidance and lacks the ability to operate independently. In contrast, agentic AI emphasizes autonomy, making it goal-driven and capable of dynamically adapting to changing environments. Unlike the prompt-dependent nature of generative AI, agentic AI is self-directed, enabling it to take the initiative and execute strategic tasks effectively. These differences highlight the complementary roles of both AI types in addressing distinct challenges.

Now let us see the impact of agentic AI on various industries. Agentic AI has had a profound impact across various industries, transforming operations and solving long-standing challenges. Autonomous logistic systems, such as those in Amazon warehouses, have significantly improved operational efficiency by 30% to 40%. In healthcare, AI-enabled surgical robots like the Da Vinci system have performed over 10 million less-invasive procedures worldwide, enhancing precision and patient outcomes. Scientific advancements have also been transformed by systems like DeepMind's AlphaFold, which successfully solved the decades-old protein folding problem. On a global scale, the World Economic Forum predicts that by 2025, AI will displace 85 million jobs while creating 97 million new ones, reshaping the labor market. And in the energy sector, AI-powered smart grids can reduce electricity waste by up to 10%, promoting greener energy solutions. Additionally, over 90 countries are investing in AI-enabled military technology to modernize their defense systems, showcasing the strategic importance of agentic AI in global security.

Now, let us see the applications of agentic AI. Agentic AI is transforming various industries by enabling systems to make autonomous decisions, adapt to changing environments, and achieve specific goals. Autonomous vehicles power self-driving cars and drones to navigate roads, avoid obstacles, and make real-time decisions, as seen with Tesla Autopilot and autonomous delivery drones. In robotics, agentic AI allows industrial, healthcare, and exploration robots to perform complex tasks independently, as demonstrated by Boston Dynamics robots used in logistics and rescue operations. Personalized virtual assistants like Google Assistant and Amazon Alexa leverage agentic AI to predict user needs, manage schedules, and execute tasks without direct commands. And in gaming, adaptive AI agents enhance the experience by creating challenging, humanlike opponents, such as AlphaGo and AI bots in real-time strategy games. In healthcare, agentic AI supports personalized treatments, accurate diagnostics, and surgical assistance, with examples including AI-driven surgical robots and systems for remote patient monitoring. These applications demonstrate the transformative potential of agentic AI across diverse domains.

Agentic AI is making a significant impact across various industries by enabling autonomy, adaptability, and efficiency in diverse applications. In finance, it powers algorithmic trading systems and fraud detection tools, optimizing financial operations such as managing investment portfolios and identifying fraudulent activities. In smart cities, AI systems manage energy consumption, optimize traffic flow, and enhance public safety, with examples like smart traffic lights adapting in real-time and autonomous energy grid optimization. In space exploration, autonomous spacecraft and planetary rovers, such as NASA's Mars rovers, perform exploration tasks independently. In education, AI-powered tutors like Carnegie Learning provide personalized instruction by adapting to individual learning styles. In military and defense, autonomous drones and surveillance systems improve situational awareness and decision-making, such as AI-driven surveillance drones in defense applications.

Now let us see the challenges and risks associated with agentic AI. While agentic AI offers tremendous potential, it also faces several challenges and risks that must be addressed to ensure its safety and ethical deployment. So, one key concern is misalignment with human goals, where AI systems may pursue objectives that conflict with human intentions due to poorly defined parameters or unintended consequences, such as an autonomous robot prioritizing efficiency over safety. Ethical questions arise regarding accountability and decision-making, demonstrated by the challenge of determining who is responsible when an autonomous vehicle causes an accident. The complexity of decision-making in agentic AI can also lead to a lack of transparency, making it difficult to understand or explain its actions, particularly in sensitive fields like healthcare or finance. Ensuring safety and reliability is another challenge, as AI systems must operate effectively in unpredictable environments, such as autonomous drones encountering extreme weather or medical failures. Additionally, agentic AI systems often require substantial computational resources, making their deployment costly, as seen in advanced robotics and self-driving cars. Security vulnerabilities pose further risks, as autonomous systems could be targeted by cyberattacks, potentially leading to harmful consequences like the manipulation of autonomous vehicles. Lastly, overdependence on AI may reduce human oversight or lead to skill degradation in critical areas, such as relying too heavily on autonomous systems for medical diagnosis without human validation. These challenges highlight the need for robust design, rigorous testing, and ethical frameworks to mitigate risks and maximize the benefits of agentic AI.

Now let's see the future of agentic AI. The future of agentic AI is set to be transformative, with advancements across various domains influencing its deployment. Future systems will exhibit increased autonomy and adaptability, enabling them to make complex decisions in real-time and operate effectively in dynamic environments without human intervention. The integration of agentic AI with advanced technologies like quantum computing, IoT, and edge computing will further enhance its capabilities, allowing for faster decision-making and real-time processing at the edge. These systems will have widespread applications in sectors such as healthcare, where they will enable autonomous medical diagnostics, personalized treatment plans, and robotic surgery; climate action, with advanced systems for environmental monitoring and response; and space exploration, where smart rovers and spacecraft will carry out missions on their own. As these technologies evolve, ethical concerns and accountability will need to be addressed, promoting the development of regulatory frameworks to ensure responsible AI usage. Additionally, agentic AI will foster human-AI collaboration, enhancing productivity and creativity in fields such as education, engineering, and research.

Imagine asking ChatGPT for a poem, and it writes one instantly. Now, think about an AI assistant planning your entire day, booking meetings, and even handling emails without your constant input. That's the difference between generative AI, which creates content, and agentic AI, which acts with autonomy, making decisions. In 2025, as AI becomes more than just a tool, understanding the shift is very critical. Are we heading towards just smarter chatbots or truly independent digital agents? Let's break it down through this video.

To truly understand the shift, let's first break down what generative AI is. Generative AI is a type of artificial intelligence designed to create content, whether it's text, images, music, or even code. Instead of making decisions or even taking action on its own, it focuses on producing outputs based on the patterns it has learned from vast amounts of data. At its core, generative AI models use deep learning techniques like transformers to generate new content that resembles human-created work. For example, ChatGPT generates humanlike text based on prompts. Midjourney and DALL-E create stunning images from simple text descriptions, and GitHub Copilot helps developers by suggesting code snippets in real-time.

Generative AI has several strengths. It enhances creativity and productivity, allowing artists, writers, and programmers to work faster and even more efficiently. It scales effortlessly, generating unlimited variations of content in just a few seconds. It also adapts responses based on user input, making interactions feel more personalized. But it also comes with a few limitations. Generative AI lacks autonomy. It doesn't think or act on its own. It only responds when prompted. It has no real decision-making abilities and cannot evaluate consequences or make even independent choices. Additionally, it can generate biased or inaccurate content based on the data that it has seen. While generative AI is powerful for creating, it cannot act independently. And that's where agentic AI comes in.

Let's explore what agentic AI is. Agentic AI goes beyond just generating content. It acts autonomously, making decisions and executing tasks without the need for constant human input. Unlike generative AI, which can only respond to prompts, agentic AI can plan, adapt, and take initiatives based on goals rather than specific instructions. At its core, agentic AI combines reasoning, memory, and decision-making to operate more like an independent agent. It doesn't just create; it analyzes, strategizes, and acts. Real-world examples include autonomous robots that navigate and complete tasks on their own; AI-driven personal assistants like those managing schedules, booking flights, and handling emails without human oversight; even self-driving cars that continuously assess their environment and make split-second driving decisions.

Agentic AI has its own strengths. It reduces the need for manual intervention, automating complex workflows. It adapts to real-world conditions, learning and improving over time. It can even handle multi-step tasks that require planning, execution, and adjustment. But it also has its own challenges. Developing truly autonomous AI requires significant advancements in reasoning and adaptability. There are certain risks, including emergent behaviors and ethical concerns around AI that makes independent decisions. And unlike generative AI, which focuses on creativity, agentic AI is limited in how well it can generate novel content. So, while generative AI creates and agentic AI acts, the real power comes when these two work together.

Let's see the key differences between generative AI and agentic AI. Generative AI and agentic AI serve different purposes, each with unique strengths and applications. The key distinction comes down to creativity versus decision-making. As previously discussed, generative AI focuses on producing content, whether it's text, image, or code. It enhances creativity by assisting writers, designers, and developers. But it lacks true autonomy. It only works when prompted and doesn't make any decisions on its own. Agentic AI, on the other hand, is designed for interactions and execution. Instead of just generating responses, it can analyze situations, make decisions, and take actions. While it may not create content like generative AI, it can manage workflows, automate tasks, and adapt to real-world conditions.

Another key difference is user dependency. Generative AI is entirely reactive, meaning it requires human input to function. It waits for prompts before generating anything. In contrast, agentic AI is proactive. It can initiate actions independently, setting reminders, optimizing schedules, or even solving problems without human intervention. The applications of these AI types also differ. Generative AI is widely used in content creation, marketing, entertainment, and software development. And agentic AI powers autonomous systems like self-driving cars, AI-powered customer service, and personal assistants that can handle complex workflows. Both AI types are transforming industries, but when they work together, they unlock even greater potential. Imagine an AI that not only generates a marketing campaign but also launches it, tracks engagement, and refines the strategy automatically. The future isn't just about choosing between generative AI and agentic AI; it's about combining the two to build truly intelligent systems.

Now that we understand the key differences between these two, let's explore the future of AI by asking, will generative AI be replaced? As AI continues to evolve, one big question arises: Will agentic AI replace generative AI? Right now, generative AI is everywhere, helping people write, design, and code faster than ever before. But it has one major limitation: it relies entirely on human input. Agentic AI, on the other hand, takes things further. It doesn't just generate; it decides, plans, and even acts. It's the next step towards true autonomous intelligence. Does that mean generative AI will be obsolete? Not necessarily. The future of AI isn't about one replacing the other; it's about coexisting. Generative AI will keep getting more creative and even sophisticated, producing even higher-quality content. Agentic AI will become even more autonomous, integrating deeper with industries like healthcare, finance, and robotics.

But this shift does come with some risks. As AI takes on decision-making power, we face new challenges: ethical concerns, unintended consequences, and the need for accountability. If an AI agent makes a bad decision, who is responsible? And how do we ensure it aligns with human values? The answer lies in balance. The real future of AI is a hybrid approach where generative AI fuels creativity and agentic AI drives intelligent action. Imagine an AI system that not only writes a research paper but also submits it to journals, responds to reviews, and refines it automatically. And this is where we are headed: not just smarter AI, but AI that truly works with us as both a creator and an agent. The question isn't whether agentic AI will replace generative AI; it's how we'll harness both to shape the future of intelligence.

Now that we have explored the differences between generative AI and agentic AI, let's move on to building an intelligent AI agent that can interact with our database using natural language. This means you can simply ask a question like, "Show me all the students who have scored above 80," and the agent will automatically convert it into an SQL query, fetch the data, and return the exact result from the database. No need to write complex SQL queries manually. Just ask, and the AI responds. Let's dive in and build this powerful system.

First, we need to set up a Conda environment to manage our project dependencies. To do this, we open the terminal and run the following command: `conda create -p venv python=3.10 -y`. So, `conda create` creates a new environment, `-p venv` specifies the environment path as `venv`, `python=3.10` installs Python version 3.10 inside the environment, and `-y` automatically confirms the installation without asking for approval. Once the process is complete, our virtual environment is ready, and we can move forward with setting up our agentic AI project.

Next, we'll create a file named `requirements.txt` where we'll list all the necessary libraries for our project. This will help us easily install dependencies in one go. Additionally, we'll create a `.env` file to securely store our Google Generative AI API key, keeping sensitive information separate from our main code. With these files in place, we ensure a well-structured and organized setup for our agentic AI project.

First, we will work with SQLite, a lightweight, self-contained database engine, to create and manage a student database. Let's break it down step by step. So, we'll create a file named `sql.py` and import the `sqlite3` module, which allows us to work with SQLite databases. We'll write `import sqlite3`. This module provides all the necessary functions to create a database, insert records, retrieve data, and manage connections.

Next, we create a connection to an SQLite database file named `student.db`. We'll write `connection = sqlite3.connect('student.db')`. If this file doesn't exist, SQLite will automatically create it. The `connection` object will allow us to interact with the database.

Now, we create a cursor object, which is used to execute SQL commands in Python. We'll write `cursor = connection.cursor()`. Think of the cursor as a tool that helps us send queries to the database and retrieve results.

Now, we define a SQL command to create a table named `student` with four columns. We'll write `table_info = """ CREATE TABLE student (name VARCHAR(25), class VARCHAR(25), section VARCHAR(25), marks INT); """`. Then we'll write `cursor.execute(table_info)`. The `name` stores the student's name (string up to 25 characters). The `class` stores the class's name, and the `section` stores the section of the student. And lastly, the `marks` stores the marks obtained as an integer. Executing this command creates the table in the database.

Next, we insert five student records into the `student` table using SQL `INSERT` statements. I've already created and inserted five values in the table. You can create as many as you can. Each `INSERT` command adds a new row with the student's name, class, section, and marks.

Now, we retrieve and display all records from the `student` table. For that, we'll have to write `print("The inserted records are:")`. In the next line, we'll write `data = cursor.execute("SELECT * FROM student")`. Then we'll write `for row in data: print(row)`. The `SELECT * FROM student` query fetches all the data from the table. The `for` loop iterates through the records and prints them one by one.

And finally, we commit our changes and close the database connection. For that, we'll write `connection.commit()` and then `connection.close()`. The `commit()` function ensures all the changes are saved in the database. The `close()` closes the connection, freeing up system resources. And that's it. We have successfully created a student database, inserted records, and retrieved them using SQLite in Python.

Now let's build an interactive Streamlit app that converts natural language questions into SQL queries using Google's Gemini model. It then retrieves data from an SQLite database and displays the result. Let's break it down step by step. But before we start, we have to activate the environment. For that, we'll write `source venv/bin/activate` (on Linux/macOS) or `.\venv\Scripts\activate` (on Windows). And here our environment is activated.

First, we'll create a file named `app.py` and load environment variables using `dotenv`. For that, we'll write `from dotenv import load_dotenv`. Next, we'll write `load_dotenv()`. It will load all environment variables. This ensures that sensitive information, such as API keys, is securely stored and accessed.

Next, we import the necessary modules. For that, we'll write `import streamlit as st`, then `import os`, then `import sqlite3`, and then `import google.generativeai as genai`. Streamlit here powers the web interface. OS helps access the environment variables. SQLite3 allows us to interact with the database, and Google Generative AI enables the conversion of natural language into SQL queries.

Now, we configure the Google Gemini API key. But before that, we'll have to create an API key through Google Studio itself. I've already generated one. You can create yours through Google Studio itself. Then we'll write `genai.configure(api_key=os.environ["GOOGLE_API_KEY"])`. This allows the app to use Gemini 1.5 Pro to generate SQL queries.

Then, we define a function to generate SQL queries from natural language input using Gemini. For that, we'll write `def get_gemini_response(question, prompt):`. Next, we'll write `model = genai.GenerativeModel('models/gemini-1.5-pro-latest')`. Then we'll write `response = model.generate_content([prompt[0], question])`. Then we'll write `return response.text`. The function initializes the Gemini model. It takes a question and a predefined prompt as input, and the AI model generates an SQL query as output.

Next, we define a function to execute SQL queries on the database and retrieve results. For that, we'll write `def read_sql_query(sql, db):`. Next, we'll write `conn = sqlite3.connect(db)`. Then `cur = conn.cursor()`. Then `cur.execute(sql)`. Then we'll write `rows = cur.fetchall()`. Then `conn.commit()`. Then `conn.close()`. Then we'll create a loop by writing `for row in rows: print(row)`. Then `return rows`. The function connects to the `student.db` database. It executes the given SQL query and it fetches all the retrieved records and prints them.

Now, we define the AI prompt that instructs Gemini on how to convert the questions into SQL queries. As you can see, I've already created a prompt for my own, and you can create yours according to how you want your model to function. If you want the prompt which I've used over here, you can just comment on the video, and I'll send it to you. This prompt ensures the Gemini AI generates SQL queries accurately without unnecessary text.

Now, we'll set the page configuration with a title and icon. For that, we'll write `st.set_page_config(page_title="SQL Query Generator - Edureka", page_icon="logo.png")`. Then we'll display the Edureka logo and header. For that, we'll write `st.image("logo.png", width=200)`. `st.markdown("## Edureka's Gemini App | Your Powered SQL Assistant")`. Next, we'll write `st.markdown("Ask any questions and I'll generate the SQL query for you.")`. The page title and the icon are set. A logo is displayed at the top, and the app's purpose is to introduce to the user. And before we import the logo, just make sure that you have the logo in your folder.

We take user input for a natural language query. For that, we'll write `question = st.text_input("Enter your query in plain English:", key="input")`. This allows users to type their questions, such as "Show all students with marks above 80." A submit button triggers the SQL generation process. And for that, we'll write `submit = st.button("Generate SQL Query")`. When clicked, the app processes the query and retrieves the result.

Now, we define what happens when the submit button is clicked. For that, we'll write `if submit:`. In the next line, `response = get_gemini_response(question, prompt)`. This is to convert the question to SQL. Then we'll print the response. Then we'll write `response = read_sql_query(response, "student.db")`. And this is to execute SQL on the database. Then we'll write `st.subheader("The response is:")`. Next, we'll include a loop `for row in response:`. Then we'll write `st.write(row)`. The user's question is converted into an SQL query using Gemini AI. The SQL query is executed on the `student.db` database, and the retrieved records are displayed on the Streamlit app. And that's it. The AI-powered Streamlit app allows users to ask natural language questions, which are automatically converted into SQL queries and executed on a student database.

Now let's open the terminal and run our Streamlit app. To do this, we simply type `streamlit run app.py` and hit enter. It's running, and as you can see, our agentic AI is up and running, ready to interact with our database. Let's test it by asking a simple question. We'll ask, "Give me the names of all the students." The AI processes our request, converts it into an SQL query, and retrieves the student names from the database. Perfect. As you can see, the response is generated.

Now, let's try another query. We'll say, "Give me the average of marks." And just like that, the AI calculates and returns the average marks. The response which is provided is 72.2. So, in this video, we successfully built an agentic AI that can understand natural language, generate SQL queries, and interact with our data seamlessly.

Amazon just dropped a major AI upgrade, Alexa, and it's unlike anything we have seen before. It's not just an update; it's a complete transformation powered by generative AI. But what exactly makes Alexa smarter, more conversational, and more capable? Well, in this video, we will break down how Amazon has leveraged state-of-the-art AI models to make Alexa a true AI assistant, how it compares to competitors like ChatGPT, and whether it's the future of voice AI.

Let's rewind a bit. Alexa started as a simple voice assistant in 2014. It could set reminders, play music, and control smart devices. But it had one major limitation: it wasn't really thinking, just following predefined rules. As AI advanced, assistants like Apple's Siri and Google Assistant improved. But Amazon saw an opportunity to turn Alexa into a true conversational AI. And that's where generative AI comes in.

Enter Alexa Plus, a brand-new AI-powered version of Alexa that understands context, remembers conversations, and sounds more natural than ever. Launched on February 26, 2025, Alexa Plus is Amazon's next-generation AI assistant designed to provide more natural conversational interactions and enhanced capabilities. This upgrade enables Alexa to perform complex tasks such as planning events, managing schedules, and controlling smart home devices more efficiently. Alexa Plus represents a significant evolution from the original Alexa, introducing several key enhancements. So, let us see what they are.

First, we have conversational abilities. Alexa Plus offers more natural and expansive interactions, understanding colloquial expressions and complex ideas, making conversations feel smoother and more intuitive. Building on that, it also takes a more proactive approach to assisting users. Unlike the original Alexa, which primarily responded to direct commands, Alexa Plus can anticipate user needs, such as suggesting earlier departures due to traffic or notifying about sales on desired items.

In addition, it has become more personalized than ever. Alexa Plus can remember user preferences, dietary restrictions, and important dates, tailoring responses and actions to individual needs, whereas the original Alexa had limited personalization capabilities. Beyond personalization, it also enhances task management. The new Alexa can handle complex tasks like making reservations, ordering groceries, and coordinating multiple services seamlessly, surpassing the more basic functionalities of the original Alexa. Not just that, it also integrates with more services than before. Alexa Plus connects with a broader range of services and devices, including Grubhub, OpenTable, Ticketmaster, and various smart home products, making it even more versatile. On top of all these improvements, it now has the ability to act independently. Agentic capabilities are a notable advancement in Alexa Plus.

Now that we have seen how Alexa Plus has improved, so let's dive into the technology behind it and understand how generative AI models and agentic AI capabilities power this next-generation assistant. Alexa is built on cutting-edge generative AI and agentic AI, leveraging powerful models and algorithms to process language, understand context, and execute tasks autonomously. So, let's break down the key technologies that make this possible.

Large Language Models (LLMs), the brain behind conversations. At the core of Alexa Plus is an advanced transformer-based language model similar to GPT-4, Claude, and Amazon's proprietary Titan model. These LLMs are trained on vast datasets, allowing Alexa to understand complex queries and respond naturally. They also maintain context across conversations, making interactions feel more fluid and generate humanlike responses, reducing robotic and repetitive phrasing. And by using techniques like reinforcement learning with human feedback, Alexa Plus continuously improves its conversational abilities based on real-world interactions.

The next technology is agentic AI, enabling proactive and autonomous actions. Beyond just responding to commands, Alexa Plus integrates agentic AI models, which allow it to act independently. Built on "do, reason, and act" models, it can plan multi-step tasks, for example, finding a restaurant, booking a table, and arranging transportation. It retrieves real-time web data to provide the latest information and executes actions across multiple apps and services without user micromanagement. This enables a fully autonomous AI assistant experience, reducing the need for manual user input.

After agentic AI, the technology that makes Alexa so versatile is neural network architectures, enabling speech and context awareness. Alexa Plus utilizes deep learning techniques such as sequence-to-sequence models for natural language generation, BERT (Bidirectional Encoder Representations from Transformers) for understanding user intent with greater accuracy, and Whisper ASR (Automatic Speech Recognition) for improved voice processing, making Alexa more responsive to different accents and speech patterns. These advancements enable highly accurate speech recognition, contextual understanding, and real-time adaptation to user behavior.

Alexa Plus integrates long-term memory storage using vector databases like Pinecone or Amazon Aurora, allowing it to remember user preferences over time, adapt to individual habits and routines for a more personalized experience, and also provide contextual reminders based on past interactions. This deep personalization is what makes Alexa Plus feel more like a true digital assistant rather than just a voice control device.

Then comes the technology that makes Alexa Plus capable of understanding and interacting with users, which is multimodal AI. Alexa Plus leverages multimodal AI, combining natural language processing for text-based queries, computer vision for Echo Show devices, enabling it to process and analyze on-screen content. Also, text-to-speech synthesis is used to generate humanlike voice responses, and this makes Alexa Plus capable of understanding and interacting with users in multiple ways, enhancing its overall functionality. By combining LLMs, agentic AI, deep learning models, and real-time data retrieval, Alexa represents a significant leap in AI-driven virtual assistance. It is no longer just a voice assistant; it is an autonomous, context-aware, and highly personalized AI companion designed to make daily life easier.

Now that we have explored the technology behind Alexa Plus, so let us see how it stacks up against other leading AI assistants. Alexa Plus enters the AI assistance space with generative AI and agentic AI, making it smarter and more proactive. But how does it compare to the top AI models available today? So, let's break it down across key aspects. So, we will compare them based on five key factors: AI power and capabilities, personalization and memory, proactive and autonomous task execution, the ecosystem and third-party integration, and finally, conversational abilities.

So, first, let us compare it with AI power and capabilities. So, how powerful is the AI behind each assistant? Alexa Plus uses Amazon Titan plus custom LLMs with generative AI and agentic AI for smart, proactive responses. Whereas ChatGPT Voice runs on GPT-4, great for deep conversation but lacks real-world task execution. Whereas Google Assistant uses Gemini AI, best for search and multimodal inputs such as text, voice, and images. And Apple Siri uses Ajax LLM, improving in language but still rule-based and limited. So, Alexa Plus leads in proactive AI, while GPT-4 dominates in conversation.

Next, let us compare it in terms of personalization and memory. So, can the assistant remember your preferences and adapt? Let us see. So, Alexa Plus has long-term memory of routines, preferences, and contextual adaptation. ChatGPT Voice has limited memory, resetting after sessions. Whereas Google Assistant remembers preferences inside Google apps but lacks deep personalization. Apple Siri has minimal memory, mostly relying on Apple's preset commands. So, Alexa Plus leads in remembering and adapting to users.

Next is proactive and autonomous task execution. So, can it handle tasks on its own? Let's see. Alexa Plus uses agentic AI for multi-step automation, for example, booking, ordering, and reminders. ChatGPT Voice assists with planning but can't perform real-world automation. Google Assistant can set reminders and retrieve information but lacks deep automation. Whereas Apple Siri is limited to commands, relying on shortcuts for basic automation. So, here, Alexa Plus is the most proactive, handling tasks automatically.

Next is the ecosystem and third-party integration. How well does it work with other devices and apps? Well, Alexa Plus is best for smart home, such as Amazon Echo, Ring, and third-party integrations. ChatGPT Voice can connect to some external tools but has no smart home control. Google Assistant has deep integration with Google apps and services. Whereas Apple Siri is limited to Apple devices with minimal third-party support. So, here again, Alexa Plus and Google Assistant lead, but Alexa has better smart home control.

And finally, conversational abilities. So, how natural and humanlike are the conversations? Alexa Plus is natural, expressive, and context-aware. ChatGPT Voice is best for deep, intelligent conversations. Google Assistant is accurate but more search-focused. Apple Siri is still command-based with limited depth. So, ChatGPT Voice is best for deep conversations, but Alexa Plus is most natural for voice interactions.

So, now let us see the future of AI assistants. Let's see what's next. So, here we have smarter AI memory: assistants will remember and personalize even better. Next, more autonomy: AI will handle complex, multi-step tasks independently. And then, more humanlike conversations: AI will feel more natural and intuitive. Next, seamless integration: AI will connect effortlessly across devices and services. Next, real-time decision-making: AI will anticipate needs and offer proactive help.

So, Alexa Plus is best for automation, memory, and smart home control. Whereas ChatGPT Voice is best for deep, intelligent conversation. Google Assistant is best for search and Google productivity. Whereas Apple Siri is best for Apple users but still limited in AI features. And Alexa Plus is not just an upgrade; it's a redefinition of AI assistance with generative AI and agentic AI for smarter, proactive help. So, what do you think? Which AI assistant is your favorite? Let me know in the comments below.

AI is no longer just responding; it's acting, planning, and automating entire workflows. Welcome to the era of agentic AI, where AI agents can write code, run businesses, and make decisions without human input. By 2030, AI automation is projected to be a $200 billion industry. And those who master agentic AI tools like AutoGPT, Devin AI, and LangChain will lead the future. Now, let's dive into the ultimate roadmap to mastering agentic AI.

So, first, let's see how you can build a strong foundation in generative AI. To truly master agentic AI, you need a strong foundation in generative AI. Understanding how AI models work, their evolution, and their impact on automation. So, start by exploring how AI has evolved from rule-based systems to advanced models like ChatGPT, AutoGPT, and Devin AI. You can check out Edureka's video on "What is Generative AI" and "Generative AI Examples" for valuable insights into the fundamentals of generative AI, its real-world applications, and how it is transforming various industries.

So, first, understand the core concepts of agentic AI, where AI can perceive, plan, and act independently to automate complex workflows. Next, learn about real-world applications such as business automation, AI-powered software engineering, and autonomous research agents. And to deepen your knowledge, familiarize yourself with the key AI models like GPT-4 Turbo, Claude AI, Gemini, and Mistral. And stay updated on multi-agent systems and self-improving AI trends. And for hands-on exploration, leverage OpenAI's API, Claude AI, and Llama 3, or experiment with different AI models on Hugging Face Spaces. You can also stay updated with AI research papers from arXiv and Hugging Face to keep up with the latest breakthroughs. Edureka's Generative AI certification and training will teach you Python programming, data science, artificial intelligence, natural language processing, and so many other updated technologies that a beginner or advanced learner is seeking. And by understanding these concepts and experimenting with these tools, you will have a strong foundation to start working with agentic AI.

Next, let's dive into programming for AI. To build and experiment with agentic AI, you need to understand the fundamentals of programming, especially in Python, which is the backbone of AI development. Start by learning Python basics, focusing on data structures, loops, functions, and object-oriented programming. Then explore essential AI and machine learning libraries like NumPy and Pandas for data manipulation, Matplotlib and Seaborn for data visualization, and TensorFlow and PyTorch for deep learning. To work with AI agents, you must also understand API interactions, as most AI tools like OpenAI's API, LangChain, and Hugging Face models require API calls. Additionally, learning automation with FastAPI, Flask, and web scraping can help you integrate AI into real-world applications. For hands-on practice, start small projects like building a chatbot, creating an AI-powered summarizer, or automating data analysis. You can also explore Edureka's Python training and certification course, designed by industry experts, where you will learn Python from scratch along with key libraries like NumPy, Pandas, Matplotlib, and Scikit-learn through hands-on projects and real-world applications. With these powerful prompting techniques and tools, you will be able to optimize AI responses and unlock the full potential of agentic AI.

To leverage agentic AI, start experimenting with cutting-edge tools that enable autonomous workflows. AutoGPT and CrewAI allow you to create multi-agent AI systems where AI agents collaborate to complete tasks. BabyAGI is perfect for automated research and decision-making, helping AI iterate on tasks dynamically. Devin AI, the first AI software engineer, showcases how AI can independently write, debug, and deploy code for hands-on learning. Build real-world projects like AI-powered automation assistants, autonomous research tools, or self-improving chatbots to see agentic AI in action. By working with these tools, you will understand how AI can move beyond just responding to acting intelligently and autonomously.

Next, explore LangChain and RAG, powerful tools that give AI the ability to retrieve real-time information, process external data, and enhance decision-making to build more powerful and context-aware AI applications. Understanding LangChain and RAG is essential. LangChain is a must-learn framework that enables seamless integration of LLMs with external data sources, allowing AI agents to interact with APIs, databases, and documents. RAG enhances AI models by providing memory and real-time knowledge retrieval, making responses more accurate and up-to-date. For hands-on learning, try building your own AI chatbot with LangChain capable of retrieving real-time information instead of relying on static training data. A great project idea to explore is an AI-powered research assistant capable of summarizing papers, fitting real-world data, and answering domain-specific questions. And to dive deeper, check out our dedicated video on LangChain and RAG, where we cover everything in detail.

Next, here are extra tips for your success. To excel in agentic AI, consistent practice and community engagement are key. So, start by pushing your AI projects to GitHub and using version control like Git to track your progress and collaborate. Join AI communities on Discord, Twitter, and Hugging Face Spaces, where you can interact with experts, stay updated on trends, and get feedback on your work. Take advantage of AI internships and open-source projects to gain real-world experience and build a strong portfolio. Also, stay updated by regularly reading AI research papers on arXiv and Google Scholar, keeping up with the latest advancements in multi-agent AI and automation. And by following these extra tips, you will accelerate your AI learning with career growth.

Think about this: Instead of you doing all your work, you have a machine to finish it for you, or it can do something which you thought was not possible. For instance, predicting the future, like predicting earthquakes, tsunamis, so that preventive measures can be taken to save lives. Chatbots, virtual personal assistants like Siri in iPhones, Google Assistant, and believe me, it is getting smarter day by day with deep learning. Self-driving cars – it will be a blessing for elderly people and disabled people who find it difficult to drive on their own, and on top of that, it can also avoid a lot of accidents that happen due to human error. Google AI Eye Doctor: So, this is a recent initiative by Google where Google is working with an Indian eye care chain to develop an AI software which can examine retina scans to identify a condition called diabetic retinopathy, which can cause blindness. AI Music Composer: Who thought that we can have an AI music composer using deep learning? And maybe in the coming years, even machines will start winning Grammys. And one of my favorites, a dream-reading machine. With so many unrealistic applications of AI and deep learning that we have seen so far, I was wondering whether we can capture dreams in the form of a video or something. And I wasn't surprised to find out that this was tried in Japan a few years back on three test subjects, and they were able to achieve close to 60% accuracy, and that is amazing. But I'm not sure whether people would want to be a test subject for this or not because it can reveal all your dreams. Great.

So, this sets the base for you, and we are ready to understand what is artificial intelligence. Artificial intelligence is nothing but the capability of a machine to imitate intelligent human behavior. AI is achieved by mimicking a human brain by understanding how it thinks, how it learns, and works while trying to solve a problem. For example, a machine playing chess, or a voice-activated software which helps you with various things on your phone, or a number plate recognition system which captures the number plate of an overspeeding car and processes it to extract the registration number and identify the owner of the car so that he can be charged. And all this wasn't very easy to implement before deep learning.

Now let's understand the various subsets of artificial intelligence. So, till now, you'd have heard a lot about artificial intelligence, machine learning, and deep learning. However, do you know the relationship between all three of them? So, deep learning is a subfield of a subfield of artificial intelligence. So, it is a subfield of machine learning, which is a subfield of artificial intelligence. So, when we look at something like AlphaGo, it is often portrayed as a big success for deep learning, but it's actually a combination of ideas from several different areas of AI and machine learning, like deep learning, reinforcement learning, self-play, etc. And the idea behind deep neural networks is not new, but it dates back to the 1950s. However, it became possible to practically implement it only when we had the new high-end resource capability. So, I hope that you have understood what is artificial intelligence. So, let's explore machine learning followed by its limitations.

So, machine learning is a subset of artificial intelligence which provides computers with the ability to learn without being explicitly programmed. In machine learning, we do not have to define all the steps or conditions like in any other programming application. However, we have to train the machine on a training dataset large enough to create a model which helps the machine to take decisions based on its learning. For example, if we have to determine the species of a flower using a machine, then first we need to train the machine using a flower dataset which contains various characteristics of different flowers along with the respective species. As you can see here in the image, we have got the sepal length, sepal width, petal length, petal width, and the species of the flower too. So, using this input dataset, the machine will create a model which can be used to classify a flower. Next, we'll pass on a set of characteristics as input to the model, and it will output the name of the flower. And this process of training a machine to create a model and use it for decision-making is called machine learning.

However, this process had some limitations. Machine learning is not capable of handling high-dimensional data, that is, where input and output are large and it is present in multiple dimensions. And handling and processing such data becomes very complex and resource-exhaustive, and this is termed as the curse of dimensionality. So, to understand this in simpler terms, let us consider a line of 100 yards, and let us assume that you dropped a coin somewhere in the line. You'll easily find the coin by simply walking on the line. A line is a single dimension.

Dimension entity. Now let's consider that you have got a square of side 100 yards each, and you dropped a coin somewhere inside the square. Now, definitely, you'll take more time to find the coin within that square. A square is a two-dimensional entity.

Now let's take it a step ahead and consider a cube of side 100 yards each, and you dropped a coin somewhere inside the cube. Now, it is even more difficult to find the coin. So, if we see that the complexity is increasing as the dimensions are increasing, and in real life, the high-dimensional data that we're talking about has got many dimensions, which makes it very, very complex to handle and process.

The high-dimensional data can be easily found in use cases like image processing, natural language processing, image translation, etc. And machine learning was not capable of solving these use cases, and hence deep learning came to the rescue. So, deep learning is capable of handling the high-dimensional data and is also efficient in focusing on the right features on its own. And this process is called feature extraction.

Now let's try and understand how deep learning works. So, in an attempt to re-engineer a human brain, deep learning studies the basic unit of a brain called a brain cell or a neuron. And inspired from a neuron, an artificial neuron or a perceptron was developed. So, if we focus on the structure of a biological neuron, it has got dendrites, and these are used to receive inputs. And these inputs are summed up inside the cell body, and using the axon, it is passed on to the next biological neuron. So, similarly, a perceptron receives multiple inputs, applies various transformations and functions, and provides an output.

As we know that our brain consists of multiple connected neurons called a neural network, we can also have a network of artificial neurons called perceptrons to form a deep neural network. Let's understand how a deep neural network looks like. So, any deep neural network will consist of three types of layers: the input layer, the hidden layer, and the output layer. So, if you see in the diagram, the first layer is the input layer, which receives all the inputs. The last layer is the output layer, which gives the desired output. And all the layers in between these layers are called hidden layers. And there can be 'n' number of hidden layers, thanks to the high-end resources available these days. And the number of hidden layers and the number of perceptrons in each layer will be entirely dependent on the use case that you're trying to solve. And there is a mechanics to decide the number of hidden layers. However, we'll not get into that in this session.

Now, since you have a picture of a deep neural network, let's try to get a high-level view of how a deep neural network solves a problem. For example, we want to perform image recognition using deep networks. So, we'll have to pass this high-dimensional data to the input layer. And to match the dimensionality of the input data, the input layer will contain multiple sub-layers of perceptrons so that it can consume the entire input. And the output received from the input layer will contain patterns and will only be able to identify the edges and images based on the contrast levels. And this output will be fed to hidden layer 1, where it will be able to identify various face features like eyes, nose, ears, etc. Now, this will be fed to hidden layer 2, where it will be able to form the entire faces and sent to the output layer to be classified and given a name. Now, think if any of these layers is missing or the neural network is not deep enough, then what will happen? Simple, we'll not be able to accurately identify the images, and this is the very reason why these use cases did not have a solution all these years prior to deep learning.

So, just to take this further, we'll try to apply a deep network on an MNIST dataset. So, the MNIST dataset consists of 60,000 training samples and 10,000 testing samples of handwritten digit images. And the task here is to train a model which can accurately identify the digit present on the image. And to solve this use case, a deep network will be created with multiple hidden layers to process all the 60,000 images pixel by pixel and finally will receive an output. So, the output will be an array of index 0 to 9, where each index corresponds to the respective digit. So, index 0 contains the probability of 0 being the digit present on the input image. Similarly, index 2, which has a value of 0.1, actually represents the probability of two being the digit present on the input image. So, if we see that the highest probability in this array is 0.8, and 8, which is present at the seventh index of the array. Hence, the number present on the image will be seven. So, this is how the handwritten image processing happens.

Let me practically execute this use case for you. So, this is my PyCharm IDE. First of all, let me show you the dataset. So, this is my MNIST dataset, and it has got four .gz files which get extracted when my program gets executed. Now, the program or the deep neural network using which I was able to create a model to process all these images and train my machine is this create_model_2.py, and it's a good lengthy program. So, I'll not be explaining the entire program for you, but let me tell you the technology or the framework with which I was able to implement this. So, I've been using TensorFlow, which is one of the open-source Google libraries for deep learning. And right here, I have imported TensorFlow, and then I'm using this MNIST data, and finally going ahead and creating a deep neural network. So, these all things are here. It is creating a deep neural network and the hidden layer that is required to process all these images. And finally, I'm creating a model and I'm saving this model with this name right here, model_2.ckpt.

Now, if I run this code, it is going to take a very long time. So, give it some time. It has extracted all the files, and it has started its training. It is at step zero now. So, in order to completely train this model, it is going to take 20,000 steps. Let me show you in the program as well. So, here it is. So, in here, I've set the steps to 20,000, but you can always configure it to a number that is 1,000, 2,000. However, you'll have to run this code for 'n' number of times so that you can achieve a particular accuracy. So, after executing for 20,000 times, what happens is a model is created with an accuracy of 92%. So, what does it mean? It means that if you pass a particular image, out of 100 images, the model, 92 predictions will be correct. 92 of the times, this model will be able to tell you the exact number that is present on the image. Let's now wait for this program to execute completely, otherwise, we have to wait for hours. So, I've already executed it once, and the model is already created. So, what is a model? A model is nothing but a set of files, and these three files along with the checkpoint files.

Now, there is another code which is predict_2.py, in which I'm restoring the model, and let me show you the line where I'm restoring it. So, it's right here: saver.restore('model_2.ckpt'). So, this is the name of my model. So, I'm restoring my entire model that was created after training of 20,000 steps on MNIST dataset, and I'm passing an image that is '7_0.png'. So, it is in this folder, in the test folder. So, I've got other images, and I've got here '7_0'. Now, this is an image. It's the name of the image. I'm not telling the program what is the number. So, I'll just stop this training now. And now we'll execute the prediction part where I'm restoring the model. And this model will tell me the number, the handwritten number that is there in the image. So, this was the image '7', and the prediction for this image is seven. So, my model was accurately able to identify or predict the number that was there on the handwritten image. So, let me change the image now, and let me execute this again, just to show you the image.

All right, there's one good question that I would like to take. So, Akil asked that, "How are you saving the weight for the neural net? Can you show us?" Sure, Akil. So, in the previous file that I executed, I showed you that I'm using an object called 'saver', and using the 'saver' object, I can save the entire model, and weights are also automatically saved along with this model in the checkpoint file. So, using checkpoint, we can actually reach the final state of the training, and then we can use the prediction model. I hope that is fine.

All right. So, if you see this '7', this is a handwritten image. This is somebody who writes '7' like this, with a strike in between. And now I'm passing a different image of '7'. It's '7_1'. So, it's different from the first one. And I'll run this code again. Now, this time the '7' is different, and my machine learning model should be able to predict that this is a '7' because people write '7' in different ways. Somebody likes '7' like this. Somebody writes '7' and makes it look like a '1', and they make the top part very small. So, there are different ways of writing '7'. So, however, a machine learning model should be capable enough to find out that as well. So, let me just close this and see the prediction. Our model was able to predict the '7' as well, and predict the value is '7' as well. So, let me execute it again for you. So, both the '7's were different, but still the prediction is correct for both of them.

So, now let us go back to our presentation. So, after the MNIST application, let me show you a few more applications of deep learning. The very first is face recognition. Let me give you an example. So, all of you are using Facebook, and you do spend some time on it. So, if you remember a few years back, when you used to upload pictures with your friends, it makes a box around a human face with a box appearing at the bottom to ask you to type the name of the person to tag him. So, it was able to identify that it was a human face. But now, it is able to auto-tag. It is not only able to detect faces but also identify who it is. And how is this possible? It is only possible using deep learning. And Facebook also has a deep learning library called Caffe2, using which they have applied all these things.

The next use case that is implemented using deep learning is Google Lens. This is one of those applications that has been recently launched for smartphones by Google. What does this app do? You just have to install it, open it, point your camera on a particular thing like this flower over here, and in real-time, image processing happens, and Google will get back with the entire details of the object, like the name of the flower, where it is found, etc., etc. So, if you point it at a building or any shop, it will tell you what kind of a store it is. If it's a restaurant, it will show you reviews, the ratings, menu, etc. So, what is happening here is that in real-time, you are able to use a deep learning net and get all the information you want, and these applications are really amazing because it directly brings deep learning to the end users or the common people. So, they can easily use the benefits of deep learning without worrying about what is happening in the background.

The next use case is machine translation. This is again a very important use case, and there is also an app in the Play Store, and this is called Translation App. So, here is an image that says "more chocolate." I don't know what it means. I don't even know which language it is. But with this app, what you can do is that you can capture the picture of the packet, and this app will first detect the text in the image, then extract the text like this, and then translate it for you. So, for example, it has detected the text, extracted it like this, and here it has translated "morg," which means dark, and then it writes back again on the image. So, what is happening in this particular use case is that first, an image is captured. Image processing takes place. Text is extracted through processing, and once we have the text, we translate it to the desired language, which is English in this case, and then again, image processing happens where we are writing the text on an image again. So, "more chocolate" means dark chocolate. So, this is a really great use case because this is a combination of multiple learning algorithms like CNN and RNN. So, these were a few more applications of deep learning. I hope that you found them interesting.

So, first thing which comes to our mind, there have been a lot of emphasis on this term called Artificial Intelligence. So, let's first try to understand that what is Artificial Intelligence on a very high level, and why we may need it in the first place for solving a problem. So, let's try to understand with an example. A person goes to a doctor and he wants to get checked that whether he got diabetes or not. And what the doctor would say is, "Okay, there are some tests which you need to get done," and based on the test results, the doctor would have a look, and from his experience, from his studies, and previous example patients he had seen, he would be able to evaluate the reports and say that the patient has diabetes or not. So, if you just take a step back and think, I said the doctor has experience. So, what do we mean by experience? The doctor has learned what are the characteristics of somebody having diabetes. Will it be possible if we can provide this experience in the form of data to a machine and let the machine take this decision whether a person has diabetes or not?

So, the experience which the doctor learned through his studies and his practice, what we are doing is we are taking customer data who have different reports and different parameters on different things like the glucose count in blood or the weight and height, and all these parameters about a human being, and based on that, we have fed it to a machine and told it what are the characteristics of a person who has diabetes. And from this, let's say we have 1 million customers' data, we have given it to a machine, and let the machine do this stuff from its experience, which comes from the data or historical data, to be precise, and do the same task which a doctor is doing. So, what we have done is, if you see from this example, what an artificial machine or artificial intelligent machine is doing is that it's learning from the historical data and trying to do the same thing which an experienced and intelligent doctor was doing. So, this kind of area or domain activities which human beings were doing, if we can make a machine intelligent enough to do the same task.

Why should we create these Artificial Intelligence-based machine systems on a very high level? There may be a lot of points, but if you just discuss a few points: human beings have limited computational power, and we guys may be good in terms of classifying things, like, you know, you can see your friend in a group photograph and easily can say who is your friend and who are others. You can easily listen to a language and comprehend what a person is doing, but human beings are not very good in doing a lot of mathematical computations. If you try doing a good amount of mathematical computations in your head, probably it will not be very easy. And second is that it's not possible for human beings to work continuously, let's say 24/7 a day for 30 days continuously. If we can make a machine do such stuff, one, they would be able to kind of do these computations very fast, and like we spent a good amount of time in discussing the GPUs, and I also mentioned that Google is talking about a TPU machine, Tensor Processing Units, which would be hugely changing the entire paradigm of computations, and machines would become more and more competitive, or even better than human beings in some of the fields.

This is the formal definition, but if I loosely translate it, it's basically that artificial intelligence machines are those machines which can do tasks which human beings can easily do. So, things like identifying what's written, let's say, in a license plate, or playing games, and I'm sure some of you would already heard that machines have defeated the Go champion and the chess players. Uh, now we have digital agents like Siri and others, which can understand what we want them to do and can take intelligent decisions from the text or from the voice itself. Basically, these are very high-level, and some of the fields where deep learning has made a great, great inroads are something like game playing, expert systems, self-driving cars, robotics, natural language processing. So, there may be different and new areas where we are implementing. All of you know that everything, every experience of human beings is getting digitalized, the kind of things you buy, the kind of things you watch, and what your preferences are, who do you like on Facebook, who you don't like, what kind of movies you like, and all these things in terms of reviews being captured online. So, once your data is going and captured online, there are systems which can analyze this data. So, given this huge data generation, as well as now we have machines which can process it and make some intelligent decisions, are available. So, that's why you will see there have been a lot of emphasis now in the last couple of years, a lot of new things are coming in. Some of you who have been reading these papers on different subjects, different architectures, would know that most of these architectures are not very old. It's a very dynamic field. Every day, in fact, on a weekly basis, you will be hearing about a new API or a new kind of architecture being developed by somebody. Most of the stuff we will be studying in these classes are not very old, like Convolutional Neural Network and Recurrent Neural Network. Some of their variants are as new as last year. If you guys follow TensorFlow closely, they introduced a library called Object Identification, Object Detection API, which TensorFlow has made available for everybody. You would be able to see it yourself that this API works. There are five different options of selecting different deep learning architectures or Convolutional Neural Networks for this API. But it's been able to identify human beings and all other 90 objects there with almost 99% accuracy. In some cases, even human beings would be finding it difficult to kind of see and predict what the object is. But this machine has gone even beyond human capability in terms of identifying objects.

Given that we have a fair understanding or very high-level understanding, we haven't got into details of Artificial Intelligence, but basically from a loose understanding that Artificial Intelligence is making decisions or machines making decisions which human beings were earlier doing the task, something like game playing and natural language and driving of cars. Let's understand how this machine learning and deep learning are related to Artificial Intelligence. Given this learning that now your machine is able to understand and learn from the data, we can solve multiple business problems with the help of this. So, let's take a very small example. It was like whether a person has diabetes or not. And I was mentioning that this kind of decision being taken by the doctor based on the reports he has got. And these reports have some numbers like the number of times a patient had a particular kind of issues, or what is the glucose count, and what is the BMI, and what the person's age is, and based on these numbers, the doctor was able to make this kind of decision. We can take the same analogy where we were trying to predict which species of flower it is. We can take information of patients and different attributes on different features of a report and the patient, and the machine would be able to learn from these datasets, and for a new customer or a new patient, it would be able to classify whether the patient has diabetes or not.

So, there are two sections of it. First one is the information about the patient and different characteristics of his health. So, from this, which is the number of times and glucose count till age, these are the information points about the patients, and the last column is the information whether the patient has diabetes or not. This kind of problem where we have some information points which explain what the situation is, and other in the last column, or the information of output is in some kind of classes. There's a specific type of machine learning problem it is, but as of now, the characteristic is that we have some information about the patient, and the last column is telling me whether the patient has diabetes or not. So, what your machine basically learns is that it learns all those rules. In the example which I was quoting, that earlier cases, people used to create these hand-coded rules to predict whether an event will happen or not. But in machine learning, your algorithm will learn from the historical data and see what are the combinations which decide whether a patient has diabetes or not. And these combinations would be of something of this type. It is only for illustration; it's not the real numbers, but it's for illustration that your machine or your machine learning algorithm is being able to identify these rules based on the historical data.

So, after learning, it has created: "If the glucose count is less than 99.5, if yes, then go to the next one. If no, the person does not have diabetes. If glucose count is greater than 166.5, yes, the person has diabetes, and if no, then there are further drill-down of rules." So, all these combinations or rules are dynamic in nature, and what I mean is that these rules would be changing if your data changes. And you can take the same model, can do the work whether a patient has diabetes or not, and you can take the same model and make it learn on a new dataset, let's say flower species, it would be able to learn the new rule from itself. So, the intuition like human beings were learning from examples, your machine learning algorithms also learn from examples.

But just to frame our problem statement, that machine learning, we know it learns by experiences and from the data, from the historical events. There are three kinds of problems which may be interested in solving. First one is called a supervised learning problem. And a supervised learning problem is basically occurs when you have some input variables and one output column. So, both the examples which we discussed till now, one was on the flower species, where we are taking data on different features of a flower, and then which species of flower it was. So, the last column is the dependent variable or output variable we are trying to predict, and all the information variables are called input features or input variables. So, input features or input variables, it's kind of interchangeably been used in different texts. Input and output, these are the two different sections of supervised learning. Why it is called supervised learning? Another take, if I need to explain it, that we have a column to guide the algorithm whether it's making the correct decisions or not. So, let's say your model says the person has diabetes, but the actual data says the person does not have diabetes. So, you have some kind of correction mechanism within your data itself which can help your model tune itself to make better predictions. So, this kind of output variable, some text, it's also been called a teacher variable. So, it's guiding the algorithm to decide those rules which I just went through.

Another type of machine learning is called unsupervised learning. And in unsupervised learning, we have only the input features, and our objective is that we should be able to identify the patterns within the data itself. So, some of you who are working in the telecom domain or marketing campaigns, you would be very much familiar with segmentation analysis or cluster analysis, where our objective is to identify coherent groups within the larger population. We take the customers as it is, the whole population, and based on different parameters and variables, we identify some of the groups of customers or products, whichever the business problem is, to identify which are similar in nature, so that we can take either a marketing campaign or develop new products for those specific groups.

Final is called reinforcement learning. A reinforcement learning is a kind of learning where the agent learns from the environment. So, it works kind of reward and penalty. You can think of a self-flying helicopter. So, you leave it in the environment, and it'll be deciding based on the wind speed and other parameters in the environment that how much it should fly, and the reward is that the fuel should be efficiently spent, and more time it should be spending in the environment.

So, supervised learning, as I said, that the objective here is that we have some input features, and input features would be holding information about different aspects of a given problem or customers, and we have an output variable which would be explaining whether the event happened or not, some kind of output variable. So, there are two types within supervised learning. One is called as regression, second is classification. And the differentiation happens only because of the type of output column. If your output column is the type of numerical values or continuous values, like numbers, so that would be a regression kind of problem, a very high intuition level. For example, you're working on a problem where you need to predict how much would be the sale of your company, given the information that how much they're spending on marketing, how many employees are working, which month it is of the year. And if you have this information, you are going to predict what is the million dollars of sales your company would be doing. So, these kind of problems where your output variable is of numerical values, then it's a regression problem. On the other hand, if the dependent variable is of categorical nature or of discrete values, it signifies that it's a classification problem. Given that it's a categorical value, your objective is that how you can put the different customers or products into different classes. So, that's why the name suggests classification problem. So, as I said, there are two kinds of supervised learning problems: one is regression, and the other one is classification.

So, let's take a use case where we need to predict the housing price of a particular locality. And we have information about these houses on different parameters, and these parameters are like these. Let's say, what is the crime rate in that area? How old is the home? The distance is how far it is from the city. This is from Boston housing data. So, this is, if I'm not wrong, percentage of black population or some variable. We have the description later on, and what is the actual price of the home, and all these features from crime to 'iset' is the information about the house. So, these are my input features, and the output feature is the price of the house. This is in million dollars. And our objective is that we should be able to fit an algorithm that it should be able to learn from all the historical homes which were sold based on all the features and what was the price it was sold for. And once the model is trained, it should be able to predict that how much should be the price for a given home.

Let's take an example. Let's say one of you is interested in buying a home in the Boston area, and you would like to know that what is the ballpark figure for a two-bedroom flat which is of some square feet and let's say 20 miles from a specific location, what should be the price. So, one way would be you go and talk to people and try to understand that what has been the average price, or if you have this kind of algorithm available which can help you understand that given these features, what was the price, and if you can create a regression model, it would be able to help you that given some features of the new home which you are interested in buying, what should be the price of it. So, as I said, there are two sections. One is the independent variables, all these information about the home, and the last variable is the dependent variable, which would be information about what was the price.

You see, this is a kind of scatter plot between the distance from the city and price for the home, keeping all other variables constant. We are not looking at the influence of other features, but we are looking if you just need to model or if you need to find a relationship between the distance from the city and what is the price. From this graph, you can make out that further the city, houses the lower the price would be, if you keep all other things constant. So, here it's like if you can identify this kind of relationship, that's called linear regression. But for a given distance from the city, you would be able to predict what should be the ballpark figure for a house. If you just have this information, not all other information which we have talked about. In a similar fashion, how it's been done is that there would be a relationship between the price of the house and all the features which we discussed. So, this is only a relationship between one of the variables, distance to the city, and the price of the home. In a similar fashion, we would be able to find a relationship between the price and all other features. So, all of the features, if you know like how old is the home, how big it is, what is the crime rate, and all. So, this kind of model is called a regression model, and it's a very basic equation of a straight line: Y = A + BX. Y here is called the dependent variable, A is intercept, and B is called slope, and X is called independent variable.

If you go deeper and try to understand what it's basically doing, this equation is trying to tell me that if I already know the relationship, if I know the value of A and B from my historical data, which is about different homes, given the value of X, I should be able to calculate Y. Y in our particular example is the price of the home. Let me try giving an example of what slope means. So, slope is the change in dependent variable if we change X, which is the independent variable, by one unit. So, let's say if I change X by one unit, how much change in Y happens? Help me understand what is the kind of relationship between X and Y. And A is the value which tells the value of Y when X is zero. And you can think of it something like that. If you put the value of X equal to zero, whatever the value is Y, then that's the intercept. But basically, from an intuition perspective, you can think that an intercept is the value which is there even though you don't have any information about X. For example, we were discussing the relationship between the distance from the city and the housing price. Even though the house is exactly in the city, then there would be some value, and even though the house is 100 miles from the city, there would still be some price. So, it helps us kind of intuitively understand the relationship. It also helps the line to understand where it starts, whether it starts from the origin or some place within your axis. And this kind of equation is called the equation of linear regression because if you see here, the power of X is one. So, that's why it's linear in nature. And it is also that we are fitting the relationship between X and only one of the variables.

Multiple regression, where what we do is that instead of finding the relationship between only one variable and the dependent variable, in most of the practical scenarios, the dependent variable Y is dependent on more than one feature. So, for an example, your house price is dependent on all these features, all these information points available, and all you want is that your regression model would be able to identify the relationship giving all the information together and then predict what is the price. The equation becomes Y = B1*X1 + B2*X2 + B3*X3. This kind of equation, where B, B1, B2, B3, and all these coefficients help us understand that what is the contribution of a single or of a given variable into your regression equation.

So, let's have a look at how you can fit this model in Python. So, if you look at the first block of the code, where we are saying `import pandas as pd`, `import numpy as np`, and `import matplotlib.pyplot as plt`. This is the convention in Python to import some of the libraries which we'll be using. So, these are the libraries which are required for running this module or this regression model. So, once we import these, all the functions available in these libraries, we can call them very easily, and we will be seeing that how you can call them.

So, once you have imported these libraries, if you see that we are loading the data called 'Boston', and this is the same dataset which we have been discussing in terms of a use case here. So, next line of code, so here we are importing the data and loading the Boston dataset. The next line of code, we are calling the pandas library because we have imported the pandas library as `pd`. Then we are calling a function called `dataframe` so that we can, you know, create a dataframe in Python, and we are creating it, `Boston data`. So, it will be creating `boss` as a dataframe. So, dataframe, on a very loose term, you can think of a kind of spreadsheet kind of format where your data is being put in rows and columns, and you can think of an Excel file kind of framework for a dataframe, though it will be different, but just for intuition purpose. And after importing, I'm calling `.head()`. So, what `.head()` does, it'll be giving the top 10 rows of my dataset. So, there are all the 13 columns. In Python, index starts from zero. You can see that index started from 0, 1, 2, 3. So, you can see what this data is. This line of command, which says `.columns`. So, `.columns` gives you all the features available in your dataset. So, these are the different feature names. So, these are the different column names for the data and the price. In the end of the code, I have written actually one line of code which can give you all the details that what the target variable is, what was the history of data, where it was recorded, and all. So, you can easily look at this. We are calling this `Boston.target` and we are calling it as `boss`. So, we are creating a column in our dataframe which was `boss` from `Boston.target`. So, there is another data vector available in the Boston dataset itself. And now, specifically, we are saying `y = this particular variable`. `y` we will be representing our dependent variable, all the features plus the Boston price. So, `boss.drop('price', axis=1)`. It means that we are dropping the price variable from the overall dataframe, and `axis=1` specifies that we are removing the column. So, `axis=0` represents the row-level operations in Python, and `axis=1` represents the column-level operations. `print` statement, we are just printing now the `x`. So, this `x` is all the input features of our dataset. So, all these columns which we will be using for predicting the housing price, and how it will be working. So, it's not an actual model, but what actually happening is, once we have created the model, it will be doing that, it'll be fitting a line which would be going through the actual dataset would be something like this, that price = some intercept term + B1 multiplied by crime + B2 multiplied by another variable, ZN, then B3 multiplied by another variable, and so on and so forth. And this intercept and B1, B2, and Bn would be the coefficients which your model would be learning from your already available data. And here we are showcasing top five values of our housing price. So, `y` is the dependent variable, and we are looking at what is the top five values.

So, this was only a very brief and very basic introduction that how do you import our data, how you can see what are the different columns. It has nothing to do with machine learning, but it is only for people who are new to Python and for people who have been out of touch in Python, if just want to brush up skills. And this line of code, if you have a look, which I'm highlighting now, it is we are using a scikit-learn model for `test_train_split`, and what it's doing is for both, because we have already `X` and `Y`, the test sizes, we are saying 33%. So, it's basically we are randomly selecting 33% of the data for, we are putting it separate in the test bucket so that we can test it later on. And this `random_state=5` means because we'll be randomly selecting, if you specify a random state, every time you run this code, you will be selecting the same set of elements from your dataset. It helps you understand that the variation, if you run the code multiple times by changing, the variation is not coming because of the selection of sample. It should be because of the different model changes you are making. This `.shape` function in Python specifies that what are the dimensions, and if I mention the first one, `X_train.shape` is giving me 339 and 13. Basically, it's telling me that there are 339 rows and 13 columns in the dataset. `X_test` there are 167 rows and 13 columns. And your `Y_test` is just 339 rows, and there's just one column, or it's just one vector. The number of rows in `X_train` and `Y_train` are same, and `X_test` and `Y_test` the number of rows are same because they have been selected for the same combination. So, same houses, we have selected the input features as well as the corresponding values of output, and same has been done for the test section.

What we are doing here, as I was saying that we have imported a library called scikit-learn, and scikit-learn has different modules for different machine learning algorithms, and linear regression is one of the modules in scikit-learn. So, we can call this scikit-learn module from linear regression called `lm`, and this `lm = LinearRegression()`. And in Python, it's called assignment variable. So, we are assigning `lm` as a linear regression module in the scikit-learn. And now what we are doing is, if you look at this line only, `lm.fit(X_train, y_train)`. So, basically, we are telling that use the linear regression module from scikit-learn and fit the model between `X_train` and `y_train`. So, basically, what we are telling the model is that you learn those coefficients for different X values given Y values in the training dataset. Basically, what the `fit` function does, it calculates the values of your intercept, B1, B2, B3 for all the features in your input features for a given Y variable.

So, once we have fit the model, and once that has been fit, we can use the same model for doing the predictions. So, let me remove it. It should be like this. So, `lm.fit(X_train, y_train)`. We have fit the model, and once there has been fit, we can use the learned model, which is `lm`, with a function called `predict(X_train)`. So, what it will be doing is, once it has learned those coefficients, B, intercept, and B1, B2, B3 for all the features and input, you can use the `predict` function for making the predictions for your training dataset, and you already have actual values as `y_train`, and then you can compare that how good your model is doing, and how you can compare it. If you look, that we have put together the same thing, we have used for the `X_test` dataset: `lm.predict(X_test)`, and again, the prediction has been done.

So, if you see here, I have put together as a dataframe `y_test` and `y_test_pred`, and the difference looks like this. For the first value, which was 37.6, and the actual value was 37, this value, then the predicted value, this actual value is this, and this difference between the actual value and the predicted value signifies that how much is the error in your dataset. So, had it been that your model is giving the same prediction as it was the actual value, you would say there is no error. Your model is 100% accurate, and all the predictions being made by the model is absolutely, you know, bang on. But normally, it doesn't happen. You end up having predictions which are a bit off from the actual value, and we measure the difference as one of the characteristics or one of the parameters to identify how correct your model is.

In this particular statistic, there are two metrics being used for identifying, but the most basic one used is called Mean Squared Error. And what Mean Squared Error is, it is basically the difference between the actual value and the predicted value by the model. And what I mean by this is that, let's say this is your predicted value, 37.6, and this is your actual value. What you do is, you take the difference of these two and then take the square of it. Why we take the square of it? Because in some values, the difference may be negative or positive, and if you sum it up, the difference may come to zero, and you may end up thinking that, "Okay, the model is doing really good stuff." To avoid it, what we do is that we take the square of it so that the difference between actual and the predicted becomes positive, and you can sum it up to showcase that how far your predicted values are from the actual value, and then you take a mean of it to showcase that what is the mean difference between the actual value and the predicted value. It can also be used for model comparisons.

Here I can show you that how it is working. That let's say you have some actual values, something like this, and let's say you fit a model, I fit a model, so there is one model prediction, one, another model is prediction two. So, what you can do is, you can take the difference between the actual and the predict, so 10 minus 2 is 2, and then you take the square of it, which is 4. 23 and 21, again 2, square of 2 is 4. Then the third one, the difference is 5, and then square of it is 25, and so on and so forth for all the values. You get the total value of sum of squared errors, and you divide it by the number of inputs, which is five, and you get the mean of the squared errors. And you do the same thing for the second one, and if you see it is very less, five. So, probably it would be able to help you understand that which model is doing a better job in terms of predicting the housing prices or any other numerical variable.

And there is another statistic which has been used for identifying how good your model is, which is called Mean Absolute Percentage Error. And that's basically the absolute difference between the actual value and the predicted value in absolute terms. And you sum it up all the values for all the entries and divide it by the number of all the value of absolute value of your actual values, and it can help you understand that what is the average percentage your predicted values are different from the actual values. So, sometime, if you see, it will be somewhere in percentages. So, what I've done is, I've taken the absolute difference between this value and the predictions, I sum it up, and divide it by the sum of my input values. So, whatever value comes in, you can say, "Okay, it's 5%." So, it'll be fair to say that your model is 5% off from the actual values, or the error term in your model is, let's say, 5% or 6%. And whatever predictions you're making from the model, you can keep a buffer of that percentage when you share it with the team. And what I mean is that, let's say your Mean Absolute Percentage Error is around 10%, and it's about the sales of a company. So, when you share this forecast, you say that my predictions are around 90% accurate, they may be actual sales maybe plus or minus 10 percentage. So, this can help you in giving this kind of variability in your predictions.

However, you have implemented code in Python itself, or the scikit-learn library, you can call `mean_squared_error` the function from sklearn, and it can help you calculate the MSE for a given model. So, basically, there was a very quick introduction to linear regression. Though there are different applications, but one thing remains common that

We are trying to predict the dependent variable, whose nature or the type of dependent variable is numerical or continuous data. Some of the applications include predicting life expectancy based on features like eating patterns, medications, disease, etc. You can predict housing prices. We have already seen the example on that. We can predict weight based on different features like sex, weight, prior information about parents, and all. And you can also predict the crop yield based on different parameters like rainfall and all. And as I said, this is a very limited use case list. I'm sure people who are working with the sales department have to make predictions on how much would be the sales. People who are working with call centers need to predict what would be the number of calls for next month. People who are working with marketing need to predict what would be the footfall in a given company or a mall. So there are different applications of regression models, but one thing is common across all these applications: the dependent variable, which we are trying to forecast, is of continuous data type.

So, let's get moving to the next agenda for logistic regression. At the time of the introduction to machine learning, we discussed there are two kinds of supervised learning techniques: one is regression, and the other one is classification. And the major differentiating factor between the two was that in regression, we had a dependent variable of continuous values, and in classification problems, we had a dependent variable of categorical types.

So, let's take an example of how we can do it. Here, let's take a use case where we have some information about some customers, and the data set looks like this: we have some customer ID or user IDs, gender of the customer or the user, his or her age, estimated salary every month (so you can think of it in any one of the currencies, either INR or dollars), and whether this user purchased an SUV or not. So, as I was saying earlier, the dependent variable here is 0 and 1. So, it's a discrete value or categorical value which we need to predict, and the features which we'll be using in the model are age and estimated salary.

Why would we need a logistic regression kind of algorithm? It would be a straightforward process that if I take "purchase" as a numerical value, 0 and 1, and I take some input features like age and estimated salary, you will be right in saying, to some extent, that this is a possibility of doing it. So, there are two major problems coming if we follow this, and some of you can help me with what may be the problems if I try using linear regression for solving this kind of problem. But one limitation I can think of is that here I'm looking for an output which can give me some kind of probability that how likely I am to buy a product or service. So, one thing, the limitation or the restriction with probability is that the probability term should be between zero and one. Zero signifies that there is no probability or there is no likelihood of an event happening, and one means that it's certain that the event will happen. There is no possibility that we can have probability values less than zero or greater than one. So, if I'm fitting a linear model, taking the "purchased" column as my dependent variable, my values, because linear regression has no such limitations, can pass these values beyond one or less than zero.

So, what I require is that I fit the model in a similar fashion like I did linear regression, the equation I used earlier. That I want that information would be coming from my features in a similar fashion. But what I want actually is that this 'y' should be mapped to values between 0 and 1. And given the limitation we have just talked about probability, that it should be between 0 and 1, it should not go beyond one and less than zero. I need to find ways if I can kind of force-fit or kind of force this 'y' value, which would be coming from this equation, and I force-fit it into values between 0 and 1.

To solve this problem, there was a function called sigmoid activation function, which would be extensively being used in our deep learning as well, at different places. But logistic regression comes from this activation function itself, which is a function that looks something like this: output value would be 1 / (1 + e^(-x)), and 'x' is not actually one of the inputs, but any value we are giving it. And if you fit any value into this particular equation, it can convert any value between minus infinity to plus infinity. It will map it to between 0 and 1. If your value of 'x', which you are putting in here, I could have selected a different value, a different name at least. But if you give a highly negative value, the output would be very, very close to zero. If this input is positive, then the output would be close to 1. If the value of 'x', the input here, becomes zero, what would be the output? 1/2, because any value to the power of 0 is equal to 1, and 1 / (1 + 1) would be equal to 1/2 or half.

So, logistic regression is nothing but an extension of your linear regression itself, with one additional fact: that you want to force-fit your output between zero and one, and for that, you are using an activation function called sigmoid activation function, or sometimes it's also been called logistic activation function, to do the same task. So, this is an intuition behind your logistic regression, where you take values of your equation from the intercept and different coefficients for your input, and you map these outputs between 0 and 1.

So, once we have understood that logistic regression is nothing but the extension of your linear regression only, with a restriction on the output being mapped between 0 and 1, we are shifting. Had we been fitting the regression equation, we would be having scenarios where the value would be going beyond one or less than zero. And to avoid this scenario, we fit in logistic regression with the help of the sigmoid activation function, which looks like this. And if you see, as I was saying, when your value of your model goes beyond, let's say this is R0, so all the values which are positive and greater than zero, the curve goes and tangential towards one. And for all the values which are less than zero, it goes towards zero. And at the place of zero, the probability is 0.5, so it's a 50% probability. If your output is very much close to zero or it's zero, it can be used for multiple scenarios. One of the examples we are taking is the example whether somebody will buy an SUV or not. But if you're trying to solve problems like somebody will say yes or no to a product or service, or whether something is true or false, or high, low, or any different categories, theories, but logistic regression can easily be put in for multiclass classification problems. And basically, if I just give you a very quick introduction on how it works, in multiclass classifications, it kind of does mapping that one class versus the rest of the other classes, and then the same analogy follows that which class a particular event would be associated with. But at the end of the day, for whichever class or category the probability is highest, the model will predict that it should belong to that particular class.

Like MSE, we have a statistic or a parameter to evaluate how good your model is doing. And that was a parameter to check what is the difference between the actual value and the predicted value. And how we were doing it, we were taking the value which was actual, subtracting the predicted value, taking the square of it, doing it for all the examples, and dividing by the number of training examples we have, and then it gives you some number. And I was also saying that this number is helpful in kind of comparing different models. So, let's say you fit a model, I fit a model, and we compare MSE for both of them. Whichever model is giving me a lesser MSE, it is kind of an indication that probably your model is doing a better job in terms of prediction than mine.

In a similar fashion, we needed a kind of a statistic to see how good your model is doing when your model is doing a classification problem. So, here there are four categories. Let's say we have only two classes: good and bad. Actual values: good and bad. And what your model is predicting: good and bad. So, four examples which belong to the good category, and your model is also predicting them as good category. So, this type of events or examples are called true positives because your model is doing correct prediction on positive examples. Another category, which is your actual value for those examples, is bad. They belong to the bad category, and your model is also predicting them as bad. These are called true negatives, and these are correct predictions because whatever the actual value is, your model is also predicting the same thing. However, there are two categories where your actual value was bad, but your model is predicting good. These kind of examples are called false positives because your model is falsely predicting them as good examples. And another category, or the last category, is called false negative, where the actual value was good, and your model is predicting bad.

So, how do you learn or how does your model say that which model is doing a good job? So, what we do is we calculate what is the percentage of values, examples have been predicted correctly. These sections in blue, true positive and true negatives, these are the examples which your model has been able to predict correctly. And these two groups, false positive and false negative, are the incorrect predictions. So, what we do is we just want to take what is the percentage of correct predictions. And this matrix is also called a confusion matrix. Let's say there are some examples out of which 65 examples were there where actually they were good category examples, and your model is also predicting them as good class, good category examples. 44 are those where they belong to the bad category, and your model is also predicting them as bad. This is 44. And eight are actual bad and prediction is good. And four are actually good and predicted bad. What you do is you sum up all the correct examples, 64 and 44, and divide it by all the examples in your data set, all correct and incorrect ones, and here you get 89%. So, all you can say is that your model is being able to predict 89% accurately. Or if you want to explain it to your business team and say that whenever I give you a prediction that 100 customers will churn, and I give you a list of 100 customers, I can say with certainty that at least 89 will churn from them with some certainty. So, because your model has given you 89% accuracy. So, that's how it's been kind of communicated to business teams that we are thinking that our model is 99% accurate, and whatever prediction we are giving, we are very, very certain. But if your model accuracy is 70% or 60%, then when you give the predictions to your business team, you say that, "Okay, though we are giving you the predictions, but we are not very certain whether it will work correctly or not." So, these accuracy percentages, in a similar fashion like we did for linear regression as MSE to identify how good the predictions are, your accuracy percentage is another metric to see how close or how correct the predictions have been.

So, now we can see the implementation of logistic regression in Python. So, first few lines, if you see, we are importing the libraries or the machine learning libraries which we require to do the data manipulations. We are importing the data, which is a CSV format, and this data is already available on your LMS. If you want to import, you can easily import from the LMS itself. Unlike the Boston data set, which we were importing from the library itself, here we have got a flat file as social_network_ads.csv, and you can call the read_csv function of the pandas library to import the data. So, you are importing the data as a dataset, and as I said, head showcases the top five rows of your data. So, here we have only five columns: one is User ID, Gender, Age, and Salary. And the last column is our dependent variable, which signifies whether a customer or a user bought the SUV or not. So, it's 0 and 1, and 1 means the person bought.

In the previous code, we used one convention of selecting X and Y. Here, we are showcasing another way of selecting that's called iloc. So, we are looking for the location, and this convention, if I go through what we are doing here, that this is the dataset, within the dataset, we are specifying the locations. This colon means that we want all the rows. And as I was saying earlier, that in Python, the index starts from zero. So, what we are saying: we want column 2 and 3. So, what we mean is that this is 0, this is 1, this is 2, and 3. So, we want as our input features 2 and 3, and the values. So, it'll be creating an array of these two columns. We could have used gender, but I will leave it to you. First, we need to create the gender as a vector of 0 and 1. So, you can create a function which will say, "Okay, if gender is equal to male, then 1, else 0," or you can create dummy variables. There are functions available in scikit-learn. So, it's an exercise for you. This is the code already available, but I would encourage that if you can also include gender information into your model.

The next line of code, why we are saying the dependent variable is all the rows and column number four. So, column number four is your purchase information, whether a customer bought the SUV or not. And again, the values to convert into a kind of list format. So, we have specified two things: that two columns is the information about the input features, or the information about the user in terms of how much money they make and what their age is, and information of why, whether a customer bought the SUV or not. And the next line, we are doing the train and test split for the same stuff to evaluate whether the predictions being made by the model on the training dataset, on which the model was learned, is still doing the correct classification on the dataset or the test dataset, which was not involved at the time of training. And this 0.25 means that we are selecting 25% of the data for test and remaining 75% for our train.

What is the correct split of train and test? Normally, it is correct to choose between something like 60-40 or 70-30 or 80-20. If your dataset is big enough, then I think having an 80-20 kind of split is good, or whatever you can try these different combinations. But as a rule of thumb, most of the time I have seen people taking something like 60-40 or 70-30 kind of distribution between actual value and the predicted value.

Now, there is one important thing for data pre-processing and this selection which I have made for doing the data pre-processing. And some of you who come from the machine learning background will already know how important it is to kind of scale your data. And what do I mean by scaling? That if you look at the dataset which we are using for input, one is the age column, and second one is the income column. Age can be somewhere between, let's say, 1 to 100 or 120, 20 at max. And your income is in like some thousands and some hundred thousands numbers. Both these values are on different scales. Scaling your data on, let's say, all the values between 0 and 1 will help me understand what is the importance of each variable. For example, if you look at a regression equation and you see those coefficient B1, B2 for all the input features, these features or these coefficients can give you a kind of indication that how important a particular variable is. But this intuition will only be correct if all my features were on the same scale. If these features like age and salary, when they are on different scales, you will not be able to compare what these coefficients really mean because they are on two different scales, your values come from. So, it is always a good idea to have all your features on the same scale.

There are multiple ways of doing it, and there are multiple types of scaling parameters. The simplest one is called Min-Max Standardization. And what does it mean is that for a given column, let's say we are talking about the age column, which is 19, 35, 26. If I need to do Min-Max Standardization, what do I mean is that I take the value, it is, let's say, 19, and the minimum value here is, let's say, 19. I have only these five values. So, how it works is that this is the formula. This is the value, or how it's being presented: X_i minus the minimum value of the column. So, let's say age divided by max of age minus minimum of range as well. But basically, what this formula will do, if I do it for all the values in the age column, it will be converting all the values between 0 and 1. And there are other ways. I also said that there are normalization processes, which is like you take the value minus the average value, divide by the standard deviation, if I call it correctly. So, whichever method we apply, all I was saying is that these values of age and salary should be brought to the same scale. So, if I'm applying this Min-Max Standardization, I'll apply it to both my columns so that both these variables are on the same scale, and I should be able to use them in my model.

And this is again a very important thing that whenever you do standardization, you will be using this process: that you fit the normalization or standard scaler on the train data set, and you use the same learned standardization from the train on the test data set. But basically, how it helps: that your dataset on which your model is being trained, it'll be converting the values between 0 and 1 based on the minimum and maximum values. If the test dataset have different minimum and maximum values, it can have different value for the same number. So, that's why the process is that we make our standardization fit on the training dataset and use the same minimum and maximum value for test normalization as well. It gives the same scale for all the values, and for model predictions, it's very helpful.

In a similar fashion like we called a linear regression object from scikit-learn in the previous example, in the exact same way, we can call a logistic regression function from scikit-learn. So, it's a scikit-learn linear model, and we are importing logistic regression. Now, um, we are fitting the logistic regression between X_train and y_train, the same way we did it earlier for the linear regression. And once it has been fitted, we can do the prediction for the test dataset. We are also doing the same thing: that we are calling the function which was classifier for the logistic regression, and we are doing it on the X_test dataset. And here, the default probability is 0.5. So, what your algorithm is doing in the backend: for all the examples wherever the probability in X_test became greater than 0.5, it was tagged as 1, and for all the examples where the probability was less than or equal to 0.5, it was tagged as 0. And now we are calling this function called confusion_matrix between y_test and y_pred. So, we are comparing that what are the values of your true positive, true negative, false positive, and false negative. So, these are the values: that these are the true positives, these are the true negatives, and these are the misclassification values. And if you want to calculate the accuracy, you can easily do it by (65 + 24) / (65 + 24 + 8 + 3). So, all we are doing is we are trying to identify what is the percentage of correctly predicted numbers. And, uh, this is the code of section. So, it can do the prediction.

If you see what it has done, basically, if you look at the section that your regression model has fit this line, and you can see it's a straight line. And that is why in some of the texts, logistic regression is also being called a linear classifier. And why it is called a linear classifier? Because it is predominantly being made for fitting a linear equation. The logistic regression equation was y = a + b1*x1 + b2*x2 and all these coefficients and respective inputs. But the highest power of your inputs were one, and you would already know that if it's a polynomial of power one, it stands for a straight line. So, that's why you can see a straight line. There are ways, some of you would argue that, you know, we can fit a non-linear line with the help of logistic regression. But you would also concede that there are some tricks which we use for creating non-linear lines through logistic regression. For example, you introduce higher-order polynomials into your model so that the separation becomes non-linear. These kind of algorithms are really helpful only in solving the problems when the objects are linearly separable. When the separation between the objects is not linearly separable, these kind of algorithms are not very helpful, and we need to identify algorithms which can fit in non-linear hypothesis or separation boundaries between different classes.

So, let's take a use case to understand what are the simple scenarios where unsupervised learning can be used and how does it really work. So, let's take an example that we have some housing data, and housing data in terms of what their locations are. And these white dots on the screen in the blue background showcase where these homes are located. And the objective of the education officer is that he needs to find a few locations where the schools can be set up, and the constraint is that students don't have to travel much. So, given this constraint in mind, the officer needs to decide the location. There may be easily we can identify if we are not using any algorithm. So, let's say if I know that I'm an officer, I need to open three schools in the locality, and I know the information where the homes are located. I can easily see, "Okay, probably this is one location." I'm just highlighting it. And the constraint I also mentioned that students don't have to travel much. What I mean is that if you open the school here, then everybody of you would say that it's not a great location for a school, given that it's far away from the population. So, this is not the correct location. And from the perspective of identifying the home, probably these three, from a human intervention or or like some of you has been given the task without any algorithm, you can decide that if you set up the school, most of the students would be traveling less to go to a school.

So, given this problem, we can easily see that we don't have a dependent variable as such which is telling us whether it's the correct location or not. All we're doing is that we have a number of locations which we need to find schools for, and then we have home locations, and based on the distance of each home, we need to identify which are the proper locations of these schools. And another thing which is coming from the same logic: that there are no predefined classes of these locations. And one more point, if you would like to add, and some of you who have done the clustering or the segmentation job in your respective works, that these numbers we say three or four or five, it's not predefined. It is most of the time given by the business that how many clusters or segments they would be looking for. Though there are statistical ways of identifying that which is the best number of clusters should be, but basically most of the time it would be coming from somebody in the business that, "Okay, I see that, let's say I was working for one of the Indian telecom companies here quite a time back, and at that time their subscriber base was around 300 million customers." And imagine that if you're trying to create segments for this big a population, and if you create three or four clusters, you can easily understand that it would be very difficult for the marketing team or any product team to design products for such a big population. So, though statistically it may look that, "Okay, four or five unique segments are there," but you end up creating lot of small, small segments, and there may be a possibility that you will be creating 20 or 30 segments for such a big population. So, my intent of saying this number, that we are trying to identify three locations within the population, has to be decided either by business or people like you who have knowledge about the data, as well as that what kind of business they are running, and what is the final usage of this segmentation exercise.

So, let's see one way of doing our selection of these school locations is like we have already doing it. If you identify that somebody looked at the homes and see the densities where the density is high and selected the home automatically. But there are algorithms also available to do the task, and I can give the name here itself: it's the K-Means algorithm. So, first, we would like to understand how does an algorithm work if it needs to identify which is the best location. So, if you're looking at my screen, let's say our objective is that we need to identify two locations first, and we have some data, and it's a scatter plot available, and we need to identify where the school should be so that the distance from home should be minimum, if that's our objective. So, how we can do it? That let's say we randomly assign two points from the existing dataset. And actually, the easiest way is that you randomly pick two numbers from your dataset itself. And then what you do is you assign these two selected points as these are your cluster centroids. So, this is the center of your selected population. And in the second step, so once you have initialized these two random points, then the next step is that you measure the distance of all the homes from the initial selected point. So, let's say you do the distance of this home from this selected point, and again from this. That for each house from these randomly initialized points, you measure the distance from the selected point or the initialized point, and any home location, and see which distance is minimum, or which distance is less in comparison to the other from the selected initial point. So, we can easily see that this distance is smaller than this distance, and this point would be assigned to this particular group.

The first initialization step is to initialize as many number of centroids as many clusters you need. And in the second step, you do the cluster assignment. And in cluster assignment, how it's been done is that you measure the distance from these initialized points and see wherever the distance is minimum, and then assign this home to that particular segment or cluster. So, this exercise has been done for each home. I'm just trying to show for a couple of them. And based on the distance, the assignment is complete. So, this color also signifies what we have done is after measuring the distance for each home from the initial points, we have assigned these points to this cluster, and these blue points to the second cluster. And then, once this assignment is complete, it moves the centroid. So, what it'll be doing is it'll be taking the center of all the selected points, and then it'll move the centroid from the previous point to the next point based on the new assignment which has already been completed. And then what's been done is the same exercise which was done earlier in terms of cluster assignment: that we measure the distance of each home from the centroid. So, distance from this centroid and this centroid, and wherever which minimum, assign it to that particular cluster. And this process has been repeated again for both the centroids. And once the distance has been measured on the improved or changed centroid, again the assignment process has been started. So, once you have moved and then measured the distance, and then assignment also changes, like it was done in the previous step, and we continue this process till the time we have reached a location or a point where this change in assignment has stopped completely. So, once we have reached this kind of place or this kind of scenario where as many time you measure the distance from the centroid to the different points, your centroid does not change. This exercise or this point is called that your model has converged. And at that point, you can say, "Okay, these all group, there is one group of these points," or "these homes." So, this is one cluster, and the second one is this cluster. So, this is how K-Means works. It has a wide variety of applications. There is a function available in scikit-learn library. You can try implementing it. The intent of showcasing you this example of unsupervised learning was that we will be having two algorithms which come from the unsupervised learning section of machine learning, and these would be your Restricted Boltzmann Machines and Autoencoders, which work on a similar methodology of unsupervised learning.

So, in a similar fashion like we started discussing in the beginning, that where should be the location of these schools, we can use a K-Means algorithm and initialize three points randomly and do this distance measure to each home and assign the homes to a cluster wherever the distance is minimum, and we continue this process of measuring the distance and assigning it to the cluster till the time these values have been converged.

The most important task for any data scientist is not to remember which library is required or what are the codes. In my understanding, the most important thing which a data scientist should remember is that once you've been given a business problem, first you should be able to understand what kind of problem it is. Whether it's a problem of supervised learning or it's an unsupervised learning. Given it's a supervised learning problem, whether it falls into the regression type or a logistic regression type. If you can make these decisions, then for implementing the algorithm, you will find a lot of help. In fact, scikit-learn would have initial codes for almost every algorithm. So, you don't have to remember lines of code and algorithms. All you should be able to do is, once the problem has been given to you, you should be able to identify what kind of problem it is.

Most of the time in unsupervised learning, and specifically in K-Means kind of models, we use this elbow method as an indicator or help you understand that what is probably a number we should start with for starting the final implementation of your model. So, let me give you an intuition on how does it work. SSE stands for Sum of Squared Errors, and what it means is actually, if I go back a little, that suppose you have identified these two clusters. So, Sum of Squared Error would be that you take the centroid and measure the distance for the points which are associated with this cluster. So, you measure the distance for each point in the orange group and sum square all the distances, and the same exercise being done for the blue points. And whatever the total number comes in after doing this exercise, you will be getting what is the total number of squared errors. And if you have two clusters, you would have some number. And just for intuition, I'm saying that this total sum of squares is coming as 100, and that's only for intuition and example I'm taking this number to help you. Let's say there was one more cluster, somebody identified here, and all these three points, though it's blue in color, but I'm saying all these three points belong to this particular segment, and rest of these points remained same as it was previously. And as we saw with two clusters, our sum of squares was coming as 100. When we have three, you can see that these points are a bit far off from this particular cluster. So, if I'll be doing it with three clusters, this distance would be a bit less, given that now I have a point which is closer to these points. And whatever error or distance these three points were adding, it would be a bit less, given the cluster was here. And let's say this distance goes down to 95. I'm just making up some numbers. So, probably what it is telling me is that sum of squares is going down, and probably I'm finding clusters which are closer or more closer to the actual data points. And as you would know that if I'll be increasing the number of clusters in the population, this distance would be going down, hopefully. And this distance can go up to zero when every point becomes a cluster itself. So, if let's say I have 20 data points there, and I assign that every point is a cluster in itself, then just measuring the distance from the point, which would be zero, and overall SSE will become zero. So, it may start from a very high number, but it will be reducing with each cluster point or cluster you will be adding to your data.

So, this line, which is sometime being called the elbow method, what it's actually showing you. So, if you had one cluster only anywhere in the population, and you do sum of the squared distances, this was the distance. When you had two clusters, this was the distance. When you had three, these were the distance. When you had four, this was the distance. But when you had five, the sum of squared error did not reduce much. So, if you see, it's like very less. And after that, even though you keep on adding different clusters, the sum of squared errors is not going down. So, as I was saying, this process or this method is kind of an indicative method, and it gives you an intuition that if I have done it, my cluster analysis with different number of clusters, and I'm measuring the sum of squared errors for a given number of clusters, and I see that after four, that the sum of squared error is not going down. It gives me an induction that probably I have found clusters which are more or less coherent, and the population is not very much away from the centroid. From the point, you can make an assumption that probably four clusters is a good idea for my given population. But as I also mentioned, that it is just an indicative process. It's a good starting point. But you need to see how the distribution of your clusters look like, whether they solve the business problem you are trying to solve or not. And if not, whether you need to further divide the clusters which your initial model has identified.

And here, we'll be taking a very quick introduction to a third type of learning, which is called reinforcement learning. What it actually is, we have seen from the two learning types, the supervised learning and unsupervised learning. The first one was that we are trying to predict some dependent variable. In the second one, we are trying to identify some kind of structure in the dataset, or if I put it into other words, that we are trying to identify some kind of coherent groups in the population. Third one is reinforcement learning, and it's basically that an object or a system learns from the environment, and there is no right or wrong answer given to the system explicitly or in the beginning itself, like in the case of supervised learning. Here, the object would be moving in the environment. Self-flying helicopters, where they fly on its own, and they take the decision that what is the wind speed and what is the pressure around it, and they correct their procedure accordingly. And the objective they need to achieve is that they need to fly for a longer period of time.

So, here we are given an example. Let's say we have a robotic dog, and somebody needs to train it to take correct decisions. And correct things would be that it's walking on the path where people need to walk, and it's not going down from the path, and if some task is being given, it's working correctly. So, there are two components of reinforcement learning, which is called reward and penalty. If the object or the system does the correct thing, it receives some reward in terms of, you know, mathematical things. Obviously, we will be providing everything in terms of mathematical numbers. And if it does the wrong thing, it receives a penalty. And based on this, it'll keep on taking its decisions. So, like a dog, if it's walking correctly, it receives some points. Like, if a ball is being thrown, if the robotic dog goes and picks it up, it's a reward point. If it doesn't do the correct thing, it receives a penalty. So, most of like all these reinforcement learning agents working in a similar fashion. Some of you who are interested in implementing it, there is an algorithm called Deep Q. It's an algorithm where you can design your own system and you can assign what are the rewards and penalty.

Similar fashion, reinforcement learning is also interacting with the space, as I mentioned. So, self-driving cars are also one of the examples which would be receiving rewards and penalties based on whether it's running on the track, taking the right turns, and moving at the correct speed, maintaining distance from other cars which are running. So, reinforcement learning has a huge implementation or requirement for self-driving cars, or some of the components of it, not all. Some of the components also in self-driving cars are supervised learning for the point that the car needs to understand what the objects are in front of it, and all other object identification.

So, what are the real limitations of machine learning? Given that we already have all three types of algorithms: supervised, unsupervised, and reinforcement learning algorithms already available. Then why we want a new architecture or new type of algorithms for our artificial intelligence systems? First and foremost is the dimensions. And when I say dimensions, it's like the type of data we get from a lot of sources. Let's say we receive images, which is grid-like image. So, where are the pixels and what is the strength of pixels in the image? Natural language processing. So, language data comes in a different length, and you know, the work is also different in the sense that suppose you need to design a machine learning algorithm which can do language translation, and if you conceptualize this idea of language translation from a machine learning algorithm perspective, your inputs become a sequence of words, and your output is also a sequence of words. And some of you who are working in machine learning algorithms, try thinking that whether we have any algorithm currently available, like logistic regression or decision tree, which can help me even fit the algorithm or fit the problem. Leave aside how good the accuracy would be and all, but these problems which come from a different type of data source, and we are trying to solve a different kind of problem, like language translation or chatbot kind of problem, where you give a sequence and it returns you a sequence. So, these kind of architectures are already not available in machine learning. So, that is one of the reasons that we need to identify some algorithms which can deal with such datasets like images and languages. And second, it can also fit different kinds of models which are not only for predicting or classifying, but also give you some kind of values like sequence. I take an example of. So, that is one of the reasons first we are looking for a different type of architecture for solving such problems.

Then, second problem which machine learning algorithms are not very good in dealing is the dimensionality. So, we would have seen with a size of, let's say, thousand variables and let's say 100,000 rows and thousand columns, probably you can still fit some of the machine learning algorithms on top of it. But given the kind of problems we are dealing with, like images, every image, let's say it's 200x200, means 200 pixels by 200 pixels, and it's a colored image, it means there are three channels. If you do the math, 200 multiplied by 200, let's say this is your image, and it's 200 by 200, because every image is kind of a matrix only. And if it's a colored image, actually colored images are being represented in the system through three channels: red, green, and blue. So, there would be three such grids, but one on top of the other. So, the number of pixels you need to have to represent your image in the system or in your algorithm would be 200x200x3. And then you calculate how many features it would be. If I, if my math is correct, it should be like 120,000 features. So, even a simple image of such small dimensions, you end up getting 120,000 features. And plus, if you're really working on a complex problem, solving in terms of, let's say, an object identification in the images, there may be five or six objects which you need to identify, and you're dealing with, let's say, 100,000 images, then your scale of data becomes so huge for any machine learning algorithm to easily handle it, and your machine learning algorithms fail in terms of getting any interesting results out of it.

So, coming to the solution part, that we need an architecture which can not only read such data in terms of images, but it is capable enough of dealing with such huge dimensions of data. So, this is the second benefit which comes from the deep learning algorithms, and we'll be discussing how do they manage such high-dimensional data when we go and talk about different architectures.

And third, and the most important reason that we will be looking for a different kind of model structure or different kind of algorithm is for identifying the features. So, in machine learning algorithms, we as data scientists spend a lot of time in kind of curating the important features. Either first, you'll be scaling the features, and after scaling, you'll be creating the interaction variables, then you'll be creating, if the separator is not very clear, then you need to introduce high-dimensional data. Let's say it's your data point, and if you see that the line you're fitting is not separating clearly, then some of you would be trying the higher-order polynomials of your input features. So, all such things which not only are difficult to, you know, come up with, there is a lot of trial and error that which kind of transformation and which kind of variable creation. So, what kind of variable will really work for a classification problem? That's the first thing. And second is, if you're working on higher-order polynomials, what is the correct order of polynomial I need to create it? And just to give you the scale of it, let's say you are dealing with only 100 features, and you need to create second-order polynomials with interaction of all these 100 features, then you'll end up getting around 5,000 features from the second-order polynomial only. If you want to get third-order polynomials, like cube variables, or or the interaction of three variables together, then these 100 variables will come around 170,000 features. So, this creation of features is very, very difficult. If we go and start creating these features on our own, and our objective, let's say, to identify a television in the image, and we have some pictures where we need to identify. Even though you have created those features manually, and some of you who are working in the field of computer science for quite some time would know that earlier we used to use features like SIFT, SIFT features, and HOG features, but these are like kind of static features for a given object. But we may argue that this television is there in this picture here, but in other picture, it can be somewhere else. So, the feature which I'm identifying, it has to be spatial indifference, that it can be anywhere in the image. And same goes for language, that if you're dealing with language data, it should be not only able to understand the meaning of a word or how does the word fit into the sentence, but should also be able to understand what is the context of each word. But these word embeddings, neural networks help you understand what are the related words to a given word, and from that, you make predictions.

So, these broad problems of machine learning algorithms: one is they are not being able to play with or deal with different types of data like images and natural language. Second is the dimensional problem, if the dimensionality goes in like 100,000 features and all. And third is this feature creation on its own. So, these are the three basic reasons that one of you or all of you would be interested in going to one of the deep learning architectures for solving such problems. And fourthly, if I may add it, that all the deep learning architectures, given that we are putting a lot of computational power in them, they end up giving you a better accuracy both for classification and regression problems. So, that that's the fourth benefit. And how does it really work? There are different stages in a deep learning and

Why they are called deep? Because it's not just input and output like we have seen in regression, that you have a Y and X is some kind of linear equation. Here we have different intermediate fields, but there are a lot of intermediary fields for doing such complex calculations. So that all these features which I mentioned, that suppose you need to identify a television, all such features get calculated at different stages, one after the other. And in the final stage, you have very, very refined features. Not only for image, we are taking the image classification, but any problem we are trying to solve through multiple stages, your model would be able to learn these intelligent features which are really important for your classification or regression or any such problem which we are trying to solve.

So these were the few benefits for deep learning, and these are actually the broad reasons that somebody would be interested in learning the deep learning algorithms. So this is the problem statement, guys. We need to figure out if the bank notes are real or fake, and for that, we'll be using artificial neural networks. And obviously, we need some sort of data in order to train our network. So let us see how the data set looks like. So over here, I've taken a screenshot of the data set with few of the rows in it. Data were extracted from images that were taken from genuine and forged bank note-like specimens. After that, wavelength transform tools were used to extract features from those images. And these are few features that I'm highlighting with my cursor. And the final column, or the last column, actually represents the label. So basically, label tells us to which class that pattern represents, whether that pattern represents a fake note or it represents a real note. Let us discuss these features and labels one by one. So the first feature, or the first column, is nothing but variance of wavelength transformed image. The second column is about skewness. The third is kurtosis of wavelength transformed image, and finally, the fourth one is entropy of the image. After that, when I talk about label, which is nothing but my last column over here, if the value is one, that means the pattern represents a real note, whereas when the value is zero, that means it represents a fake note.

So guys, let's move forward and we'll see what are the various steps involved in order to implement this use case. So over here, we'll first begin by reading the data set that we have. We'll define features and labels. After that, we are going to encode the dependent variable. And what is a dependent variable? It is nothing but your label. Then we are going to divide the data set into two parts: one for training, another for testing. After that, we'll use TensorFlow data structures for holding features, labels, etc. And TensorFlow is nothing but a Python library that is used in order to implement deep learning models, or you can say neural networks. Then we'll write the code in order to implement the model. And once this is done, we will train our model on the training data. We'll calculate the error. The error is nothing but your difference between the model output and the actual output, and we'll try to reduce this error. And once this error becomes minimum, we'll make predictions on the test data and we'll calculate the final accuracy.

So guys, let me quickly open my PyCharm and I'll show you how the output looks like. So this is my PyCharm, guys. Over here, I've already written the code in order to execute the use case. I'll go ahead and run this and I'll show you the output. So over here, as you can see, with every iteration, the accuracy is increasing. So let me just stop it right here. All right. Till now, any questions? Any doubts with respect to what is our use case? What is the data set about? So we'll move forward and we'll understand why we need neural networks. So in order to understand why we need neural networks, we are going to compare the approach before and after neural networks, and we'll see what were the various problems that were there before neural networks.

So earlier, conventional computers use an algorithmic approach. That is, the computer follows a set of instructions in order to solve a problem. And unless the specific steps that the computer needs to follow are known, the computer cannot solve the problem. So obviously, we need a person who actually knows how to solve that problem, and he or she can provide the instructions to the computer as to how to solve that particular problem, right? So we first should know the answer to that problem, or we should know how to overcome that challenge or problem which is there in front of us. Then only we can provide instructions to the computer. So this restricts the problem-solving capability of conventional computers to problems that we already understand and know how to solve. But what about those problems whose answer we have no clue of? So that's where our traditional approach was a failure. So that's why neural networks were introduced.

Now let us see what was the scenario after neural networks. So neural networks basically process information in a similar way the human brain does, and these networks, they actually learn from examples. You cannot program them to perform a specific task. They will learn from their examples, from their experience. So you don't need to provide all the instructions to perform a specific task, and your network will learn on its own with its own experience. All right. So this is what basically neural network does. So even if you don't know how to solve a problem, you can train your network in such a way that with experience, it can actually learn how to solve the problem. So that was a major reason why neural networks came into existence.

So these neural networks are basically inspired by neurons, which are nothing but your brain cells. And the exact working of the human brain is still a mystery, though. So as I've told you earlier as well, that neural networks work like a human brain, and so the name. And similar to a newborn human baby, as he or she learns from his or her experience, we want a network to do that as well, but we want it to do very quickly. So here's a diagram of a neuron. Basically, a biological neuron receives input from other sources, combines them in some way, performs a generally non-linear operation on the result, and then outputs the final result. So here, if you notice these dendrites, these dendrites will receive signals from the other neurons. Then what will happen? It will transfer it to the cell body. The cell body will perform some function. It can be summation, it can be multiplication. So after performing that summation on the set of inputs, via axon, it is transferred to the next neuron.

Now let's understand what exactly are artificial neural networks. It is basically a computing system that is designed to simulate the way the human brain analyzes and processes information. Artificial neural networks have self-learning capabilities that enable it to produce better results as more data becomes available. So if you train your network on more data, it'll be more accurate. So these neural networks, they actually learn by example. And you can configure your neural network for specific applications. It can be pattern recognition, or it can be data classification, anything like that. All right. So because of neural networks, we see a lot of new technology has evolved, from translating web pages to other languages, to having a virtual assistant to order groceries online, to conversing with chatbots. All of these things are possible because of neural networks. So in a nutshell, if I need to tell you, artificial neural network is nothing but a network of various artificial neurons. All right.

So let me show you the importance of neural network with two scenarios: before and after neural network. So over here, we have a machine, and we have trained this machine on four types of dogs, as you can see where I'm highlighting with my cursor. And once the training is done, we provide a random image to this particular machine which has a dog, but this dog is not like the other dogs on which we have trained our system on. So without neural networks, our machine cannot identify that dog in the picture, as you can see it over here. Basically, our machine will be confused. It cannot figure out where the dog is. Now, when I talk about neural networks, even if we have not trained our machine on this specific dog, but still it can identify certain features of the dogs that we have trained on, and it can match those features with the dog that is there in this particular image, and it can identify that dog. So this happens all because of neural networks. So this is just an example to show you how important are neural networks.

Now, I know you all must be thinking, how neural networks work. So for that, we'll move forward and understand how it actually works. So over here, I'll begin by first explaining a single artificial neuron that is called as perceptron. So this is an example of a perceptron. Over here, we have multiple inputs: X1, X2, till Xn, and we have corresponding weights as well: W1 for X1, W2 for X2, similarly Wn for Xn. Then what happens? We calculate the weighted sum of these inputs. And after doing that, we pass it through an activation function. This activation function is nothing but it provides a threshold value. So above that value, my neuron will fire, else it won't fire. So this is basically an artificial neuron. So when I talk about a neural network, it involves a lot of these artificial neurons with their own activation function and their processing element.

Now we'll move forward and we'll actually understand various modes of this perceptron or single artificial neuron. So there are two modes in a perceptron: one is training, another is using mode. In training mode, the neuron can be trained to fire for particular input patterns. Which means that we'll actually train our neuron to fire on certain set of inputs and to not fire on the other set of inputs. That's what basically training mode is. When I talk about using mode, it means that when a dot input pattern is detected at the input, its associated output becomes the current output. Which means that once the training is done, and we provide an input on which the neuron has been trained on, so it'll detect the input and we'll provide the associated output. So that's what basically using mode is. So first, you need to train it, then only you can use your perceptron or your network. So these were the two modes, guys.

Next up, we'll understand what are the various activation functions available. So these are the three activation functions. Although there are many more, but I've listed down three: step function. So over here, the moment your input is greater than this particular value, your neuron will fire, else it won't. Similarly for sigmoid and sine function as well. So these are three activation functions. There are many more that I've told you earlier as well. So these are the three majorly used activation functions.

Next up, what we are going to do, we are going to understand how a neuron learns from its experience. So I'll give you a very good analogy in order to understand that. And later on, when we talk about neural networks, or you can say multiple neurons in a network, I'll explain you the math behind it. I'll explain you the math behind learning, how it actually happens. So right now, I'll explain you with an analogy. And guys, trust me, that analogy is pretty interesting. So I know all of you must have guessed it. So these are two beer mugs, and all of you who love beer can actually relate to this analogy a lot. And I know most of you actually love beer. So that's why I've chosen this particular analogy so that all of you can relate to it. All right, jokes apart. So fine, guys. So there's a beer festival happening near your house, and you want to badly go there. But your decision actually depends on three factors: first is, how is the weather, whether it is good or bad. Second is, your wife or husband is going with you or not. And the third is, any public transport is available. So on these three factors, your decision will depend whether you'll go or not. So we'll consider these three factors as inputs to our perceptron, and we'll consider our decision of going or not going to the beer festival as our output. So let us move forward with that.

So we'll move forward and we'll see what are the various inputs that I'm talking about. So the first input is, how is the weather? We'll consider it as X1. So when weather is good, it'll be one, and when it is bad, it'll be zero. Similarly, your wife is going with you or not. So that'll be your X2. If she is going, then it's one. If she's not going, then it's zero. Similarly for public transport, if it is available, then it is one, else it is zero. So these are the three inputs that I'm talking about. Let's see the output. So output will be one when you're going to the beer festival, and output will be zero when you want to relax at home. You want to have beer at home only. You don't want to go outside. So these are the two outputs: whether you are going or you're not going.

Now, what a human brain does over here? Okay, fine. I need to go to the beer festival. But there are three things that I need to consider. But will I give importance to all these factors equally? Definitely not. There'll be certain factors which will be of higher priority for me. I'll focus on those factors more, whereas few factors won't affect that much to me. All right? So let's prioritize our inputs or factors. So here, our most important factor is weather. So if weather is good, I love beer so much that I don't care even if my wife is going with me or not, or if there is a public transport available. So I love beer that much that if weather is good, that definitely I'm going there. That means when X1 is high, output will be definitely high. So how we do that? How we actually prioritize our factors, or how we actually give importance more to a particular input and less to another input in a perceptron or in a neuron? So we do that by using weights. So we assign high weights to the more important factors or more important inputs, and we assign low weights to those particular inputs which are not that important for us. So let's assign weights, guys. So weight W is associated with input X1, W2 with X2, and similarly W3 with X3. Now, as I've told you earlier as well, that weather is a very important factor. So I'll assign a pretty high weight to weather and I'll keep it as six. Similarly, W2 and W3 are not that important. So I'll keep it as two, two. After that, I've defined a threshold value as five, which means that when the weighted sum of my input is greater than five, then only my neuron will fire, or you can say, then only I'll be going to the beer festival. All right.

So I'll use my pen and we'll see what happens when weather is good. So when weather is good, our X1 is one. Our weight is six. We'll multiply it with six. Then, if my wife decides that she is going to stay at home, and she will probably be busy with cooking, and she doesn't want to drink beer with me, so she's not coming. So that input becomes zero. Zero into two will actually make no difference because it'll be zero only. Then again, there is no public transport available also. Then also this will be zero into two. So what output I get here? I get here as six. And notice the threshold value, it is five. So definitely, six is greater than five. That means my output will be one, or you can say my neuron will fire, or I'll actually go to the beer festival. So even if these two inputs are zero for me, that means my wife is not willing to go with me, and there is no public transport available, but weather is good, which has very high weight value, and it actually matters a lot to me. So if that is high, it doesn't really matter whether the two inputs are high or not, I will definitely go to the beer festival.

All right. Now I'll explain you a different scenario. So over here, our threshold was five, but what if I change this threshold to three? So in that scenario, even if my weather is not good, I'll give it a zero. So zero into six. But my wife and public transport both are available. All right. So one into two plus one into two, which is equal to four, and it is definitely greater than three. Then also my output will be one. That means I will definitely go to the beer festival, even if weather is bad, and my neuron will fire. So these are the two scenarios that I have discussed with you. All right. So there can be many other ways in which you can actually assign weight to your problem or to your learning algorithm. So these are the two ways in which you can assign weights and prioritize your inputs or factors on which your output will depend. So obviously, in real life, all the inputs or all the factors are not as important for you. So you actually prioritize them, and how you do that in perceptron? You provide high weight to it. This is just an analogy so that you can relate to a perceptron to a real life. We'll actually discuss the math behind it later in the session as to how a network or a neuron learns. All right. So how the weights are actually updated, and how the output is changing, all those things we'll be discussing later in the session. But my aim is to make you understand that you can actually relate to a real-life problem with that of a perceptron. All right? And in real life, problems are not that easy. They are very, very complex problems that we actually face. So in order to solve those problems, a single neuron is definitely not enough. So we need networks of neurons, and that's where artificial neural network, or you can say multi-layer perceptron, comes into the picture.

Now let us discuss that multi-layer perceptron or artificial neural network. So this is how an artificial neural network actually looks like. So over here, we have multiple neurons present in different layers. The first layer is always your input layer. This is where you actually feed in all of your inputs. Then we have the first hidden layer. Then we have the second hidden layer, and then we have the output layer. Although the number of hidden layers depend on your application, on what are you working, what is your problem. So that actually determines how many hidden layers you'll have. So let me explain you what is actually happening here. So you provide some input to the first layer, which is nothing but your input layer. You provide inputs to these neurons. All right? And after some function, the output of these neurons will become the input to the next layer, which is nothing but your hidden layer one. Then these hidden layers also have various neurons. These neurons will have different activation functions. So they'll perform their own function on the inputs that it receives from the previous layer, and then the output of this layer will be the input to the next hidden layer, which is hidden layer two. Similarly, the output of this hidden layer will be the input to the output layer, and finally, we get the output. So this is how basically an artificial neural network looks like.

Now let me explain you this with an example. So over here, I'll take an example of image recognition using neural networks. So over here, what happens? We feed in a lot of images to our input layer. Now, this input layer will actually detect the patterns of local contrast. And then we'll feed that to the next layer, which is hidden layer one. So in this hidden layer one, the face features will be recognized. We'll recognize eyes, nose, ears, things like that. And then that will be again fed as input to the next hidden layer. And in this hidden layer, we'll assemble those features and we'll try to make a face. And then we'll get the output that is the face will be recognized properly. So if you notice here, with every layer, we are trying to get a more abstract version or the generalized version of the input. So this is now basically an artificial neural network, how it works. All right. And there's a lot of training and learning which is involved that I'll show you now.

Training a neural network. So how we actually train our neural network. So basically, the most common algorithm for training a network is called backpropagation. So what happens in backpropagation? After the weighted sum of inputs and passing through an activation function and getting the output, we compare that output to the actual output that we already know. We figure out how much is the difference. We calculate the error. And based on that error, what we do, we propagate backwards and we'll see what happens when we change the weight, will the error decrease or will it increase? And if it increases, when it increases, by increasing the value of the variables or by decreasing the value of variables. So we kind of calculate all those things and we update our variables in such a way that our error becomes minimum. And it takes a lot of iterations. Trust me, guys, it takes a lot of iterations. We get output a lot of times, and then we compare it with the model, with the actual output. Then again, we propagate backwards. We change the variables. Then again, we calculate the output. We compare it again with the desired output or the actual output. Then again, we propagate backwards. So this process keeps on repeating until we get the minimum value. All right.

So there's an example that is there in front of your screen. Don't be scared of the terms that I used. I'll actually explain you with an example. So this is the example over here. We have zero, one, and two as inputs. And our desired output, or the output that we already know, is zero, one, and four. All right. So over here, we can actually figure out that the desired output is nothing but twice of your input. But I'm training a computer to do that, right? The computer is not a human. So what happens? I actually initialize my weight. I keep the value as three. So the model output will be three into zero is zero. Three into one is three. Three into two is six. Now, obviously, it is not equal to your desired output. So we check the error. Now, the error that we have got here is zero, one, and two, which is nothing but your difference. So zero minus zero is zero. Three minus two is one. Six minus four is two. So this is called an absolute error. After squaring this error, we get square error, which is nothing but zero, one, and four. All right.

So now what we need to do, we need to update the variables. We have seen that the output that we got is actually different from the desired output. So we need to update the value of the weight. So instead of three, our computer makes it as four. After making the value as four, we get the model output as zero, four, and eight. And then we saw that the error has actually increased. Instead of decreasing, the error has increased. So after updating the variable, the error has increased. So you can see that square error is now zero, four, and sixteen. And earlier it was zero, one, and four. That means we cannot increase the weight value right now. But if we decrease that, make it as two, we get the output which is actually equal to the desired output. But is it always the case that we need to only decrease the weight? Definitely not. So in this particular scenario, whenever I'm increasing the weight, error is increasing, and when I'm decreasing the weight, error is decreasing. But as I've told you earlier as well, this is not the case every time. Sometimes you need to increase the weight as well. So how we determine that? All right, fine, guys. This is how basically a computer decides whether it has to increase the weight or decrease the weight. So what happens here? This is a graph of square error versus weight. So over here, what happens? Suppose your square error is somewhere here, and your computer, it starts increasing the weight in order to reduce the square error, and it notices that whenever it increases the weight, square error is actually decreasing. So it'll keep on increasing until the square error reaches a minimum value, and after that, when it tries to still increase the weight, the square error will increase. So at that time, our network will recognize that whenever it is increasing the weight after this point, error is increasing. So therefore, it will stop right there, and that will be our weight value. Similarly, there can be one more scenario. Suppose if we increase the weight, but then also the square error is increasing. So at that time, we cannot increase the weight. At that time, computer will realize, okay, fine, whenever I'm increasing the weight, the square error is increasing. So it'll go in the opposite direction. So it'll start decreasing the weight, and it'll keep on doing that until the square error becomes minimum. And the moment it decreases more, the square error again increases. So our network will know that whenever it decreases the weight value, the square error is increasing. So that point will be our final weight value. So guys, this is what basically backpropagation in a nutshell is. If you have any questions or doubts, you can go ahead and ask me. All right, fine. We have no doubts here. Fine.

So we'll move forward and now is the correct time to understand how to implement the use case that I was talking about at the beginning. That is, how to determine whether a note is fake or real. So for that, I'll open my PyCharm. This is my PyCharm again, guys. Uh, let me just close this. All right. So this is the code that I've written in order to implement the use case. So over here, what we do, we import the first important libraries which are required. Matplotlib is used for visualization. TensorFlow, we know, in order to implement the neural networks. NumPy for arrays. Pandas for reading the data set. Similarly, scikit-learn for label encoding, as well as for shuffling, and also to split the data set into training and testing tasks. All right, fine, guys.

So we'll begin by first reading the data set as I've told you earlier as well when I was explaining the steps. So what I'll do, I'll use pandas in order to read the CSV file which has the data set. After that, I'll define features and labels. So X will be my features, and Y will contain my labels. So basically, X includes all the columns apart from the last column, which is the fifth one. And because the indexing starts from zero, that's why we have written zero till fourth. So it won't include the fourth column. All right. And so our last column will actually be our label. Then what we need to do, we need to encode the dependent variable. So the dependent variable, as I've told you earlier as well, is nothing but your label. So I've discussed encoding in TensorFlow tutorial. You can go through it, and you can actually get to know why and how we do that. Then what we have done, we have read the data set. Then what we need to do is to split our data set into training and testing. And these are all optional steps. You can print the shape of your training and test data. If you don't want to do it, you're still fine.

Then we have defined learning rate. So learning rate is actually the steps in which the weights will be updated. All right. So that is what basically learning rate is. Then when we talk about epoch, means iterations. Then we have defined cost history, that will be an empty NumPy array, and its shape will be one, and it'll include the float type objects. Then we have defined N_dim, which is nothing but your X shape of axis one, which means your column. So we'll print that. After that, we have defined the number of classes. So there can be only two classes: whether the note can be fake or it can be real. And this model path, I've given in order to save my model. So I've just given a path where I need to save it. So I'll just save it here only, in the current working directory.

Now is the time to actually define our neural network. So we'll first make sure that we have defined the important parameters like hidden layers, number of neurons in hidden layers. So I'll take 10 neurons in every hidden layer, and I'm taking four layers like that. Then X will be my placeholder, and the shape of this particular placeholder is none, N_dim. N_dim value, I'll get it from here, and none can be any value. I'll define one variable W, and I'll initialize it with zeros, and this will be the shape of my weight. Similarly for bias as well, this will be the particular shape. And there will be one more placeholder, Y_dash, which will actually be used in order to provide us with the actual output of the model. There'll be one model output, and there'll be one actual output, which we use in order to calculate the difference, right? So we'll feed in the actual values of the labels in this particular placeholder Y_dash. And now we'll define the model. So over here, we have named the function as multi_layer_perceptron. And in it, we'll first define the first layer. So the first hidden layer, and we are going to name it as layer_1, which will be nothing but the matrix multiplication of X and weights of h1, that is the hidden layer 1, and that'll be added to your biases b1. After that, we'll pass it through a sigmoid activation function. Similarly, in layer 2 as well, matrix multiplication of layer 1 and weights of h2. So if you can notice, layer 1 was the network layer just before the layer 2, right? So the output of this layer 1 will become input to the layer 2. And that's why we have written layer_1. It'll be multiplied by weights h2, and then we'll add it with the bias. Similarly for this particular hidden layer as well, and this particular layer as well. But over here, we are going to use ReLU activation function instead of sigmoid.

Then we are going to define the weights and biases. So this is how we basically define weights. This is how we basically define weights. So weights_h1 will be a variable which will be a truncated normal with the shape of N_dim and N_hidden_1. So these are nothing but your shapes. All right. And after that, what we have done, we have defined biases as well. Then we need to initialize all the variables. So all these things actually I've discussed in brief when I was talking about TensorFlow. So you can go through TensorFlow tutorial at any point of time if you have any question. We have discussed everything there. Since in TensorFlow, we need to initialize the variables before we use it. So that's how we do it. We first initialize it, and then we need to run it. That's when your variables will be initialized. After that, we are going to create a saver object, and then finally, I'm going to call my model. And then comes the part where the training happens. Cost function. Cost function is nothing but you can say an error that will be calculated between the actual output and the model output. All right. So Y is nothing but our model output, and Y_dash is nothing but actual output, or the output that we already know. All right. And then we are going to use a gradient descent optimizer to reduce error. Then we are going to create a session object as well. And finally, what we are going to do, we are going to run the session. So this is how we basically do that. For every epoch, we will be calculating the change in the error as well as the accuracy that comes after every epoch on the training data. After we have calculated the accuracy on the training data, we're going to plot it for every epoch, how the accuracy is. And after plotting that, we're going to print the final accuracy, which will be on our test data. So using the same model, we'll make predictions on the test data, and after that, we are going to print the final accuracy and the mean squared error.

So let's go ahead and execute this, guys. All right. So training is done. And this is the graph we have got for accuracy versus epoch. This is accuracy. Y-axis represents accuracy, whereas this is epoch. We have taken 100 epochs, and our accuracy has reached somewhere around 99%. So with every epoch, it is actually increasing, apart from a couple of instances, it is actually keep on increasing. So the more data you train your model on, it'll be more accurate. Let me just close it. So now the model has also been saved where I wanted it to be. This is my final test accuracy, and this is the mean squared error. All right. So these are the files that will appear once you save your model. These are the four files that I've highlighted.

Now what we need to do is restore this particular model. And I've explained this in detail how to restore a model that you have already saved. So over here, what I'll do, I'll take some random range. I've taken it actually from 754 to 768. So all the values in the row of 754 and 768 will be fed to our model, and our model will make predictions on that. So let us go ahead and run this. So when I'm restoring my model, it seems that my model is 100% accurate for the values that I have fed in. So whatever values that I have actually given as input to my model, it has correctly identified its class, whether it's a fake note or a real note, because zero stands for fake note, and one stands for real note. Okay. So original class is nothing but which is there in my data set. So it is zero already, and what prediction my model has made is zero, that means it is fake. So accuracy becomes 100%. Similarly for other values as well. Fine, guys. So this is how we basically implement the use case that we saw in the beginning.

So in the slide, you can notice that I've listed out only two applications. Although there are many more. So neural networks in medicine. Artificial neural networks are currently a very hot research area in medicine, and it is believed that they will receive extensive application to biomedical systems in the next few years. And currently, the research is mostly on modeling parts of the human body and recognizing diseases from various scans. For example, it can be cardiograms, CAT scans, ultrasonic scans, etc. And currently, the research is going mostly on two major areas: first is modeling and diagnosing the cardiovascular system. So neural networks are used experimentally to model the human cardiovascular system. Diagnosis can be achieved by building a model of the cardiovascular system of an individual and comparing it with the real-time physiological measurements taken from the patient. And trust me, guys, if this routine is carried out regularly, potential harmful medical conditions can be detected at an early stage, and thus make the process of combating disease much easier. Apart from that, it is currently being used in electronic noses as well. Electronic noses have several potential applications in telemedicine.

Now, let me just give you an introduction to telemedicine. Telemedicine is a practice of medicine over long distance via a communication link. So what the electronic noses will do, they would identify odors in the remote surgical environment. These identified odors would then be electronically transmitted to another site where an odor generation system would recreate them. Because the sense of smell can be an important sense to the surgeon. Tele-smell would enhance tele-presence surgery. So these are the two ways in which you can use it in medicine. You can use it in business as well, guys. So business is basically a diverted field with several general areas of specialization, such as accounting or financial analysis. Almost any neural network application would fit into one business area or financial analysis. Now, there is some potential for using neural networks for business purposes, including resource allocation and scheduling. I've listed down two major areas where it can be used: one is marketing. So there is a marketing application which has been integrated with the neural network system. The airline marketing tactician is a computer system made of various intelligent technologies, including expert systems. A feed-forward neural network is integrated with the AMT, which is nothing but airline marketing tactician, and was trained using backpropagation to assist the marketing control of airline seat allocation. So it has wide applications in marketing as well.

Now, the second area is credit evaluation. Now I'll give you an example here. The HNC company has developed several neural network applications, and one of them is a credit scoring system, which increases the profitability of existing models up to 27%. So these are few applications that I'm telling you, guys. Neural network is actually the future. People are talking about neural networks everywhere, and especially after the introduction of GPUs and the amount of data that we have now, neural network is actually spreading like plague right now.

So what is a transformer? Transformers operate on a concept called sequence-to-sequence learning. Essentially, they take a sequence of tokens as an input and predict the next token in the output. A great example of this is language translation. Imagine inputting "good morning" in English, and the transformer processes this and outputs the translation in languages like Japanese, Korean, or German. The key is how it efficiently processes the relationship between words. Since we know what a transformer is, let's dig a bit deep about them. A transformer has two primary components: encoder and decoder. Encoder identifies relationships between parts of the input sequence, whereas the decoder uses these relationships to generate the output sequence. This division is what allows transformers to handle tasks like text translation or summarization with remarkable accuracy.

Now that we have the idea of transformers, let's discuss how they evolved. Before transformers, there were other neural networks like RNN, recurrent neural networks, invented by David Rumelhart in 1986. However, RNNs faced significant challenges. They would forget early parts of the sequence as they processed longer ones and couldn't handle dependencies efficiently. Additionally, RNNs relied on recurrence, which made them inefficient and incapable of parallelization. Then came Long Short-Term Memory, introduced by Hochreiter and Schmidhuber in 1997. Long Short-Term Memory improved by remembering sequences for a longer duration and addressing some of the memory issues in RNNs. However, they were slow to train and difficult to manage at scale. Finally, transformers transformed neural networks. First introduced in the landmark paper, "Attention Is All You Need," transformers addressed all the problems faced by RNNs and LSTMs. They used a completely attention-based mechanism, eliminating reliance on recurrence. This made transformers capable of remembering context efficiently, training faster, and being parallelized, enabling multitasking and significantly speeding up processes.

Now, let's discuss on the attention mechanism. Think about the sentence: "This cat wants to jump on the box." The attention mechanism identifies the most relevant parts of this sentence, like "cat," "jump," and "box," and focuses on these elements while processing the data. Now that we know how transformers have evolved, now let's discuss their architecture. A transformer consists of two main components: an encoder and a decoder, each typically consisting of six layers. Inside the encoder, there is one attention layer and one feed-forward layer. While the decoder contains two attention layers and one feed-forward layer. The magic of parallelism comes from how data is fed into the network. In the attention layer, all the words are processed simultaneously, with each word forming combinations with others in the sentence. This allows the model to capture relationships and context efficiently. After processing in the attention layer, the data is sent to the feed-forward layer, where it is learned layer by layer. The input to the encoder and decoder are the raw input embeddings, which are numerical representations of words. On top of these embeddings, positional encodings are added to help the model understand the position and the order of each word in the sequence. If we simplify embeddings, they are essentially vector representations of words in an N-dimensional space. At the top of the architecture, there are two layers of output probabilities, converting the final output into a form that humans can understand. These inputs are represented as vectors with their length corresponding to the size of the vocabulary.

Now, what truly makes transformers unique is the inclusion of normalization layers, which normalize the output from sub-layers. Additionally, skip connections, the dark arrows in the architecture, forward critical information that bypasses self-attention or feed-forward layers directly to the normalization layers. This ensures the model does not forget important details and effectively passes vital information further into the network.

Now moving forward, let's discuss why transformers are important. Transformers are vital because they utilize semi-supervised learning. They are trained on massive unlabeled datasets, enabling them to generalize across a wide range of tasks. Unlike older models, transformers don't need to process data sequentially. Their attention mechanism allows them to focus on the most relevant context, which significantly speeds up training. Transformers revolutionize data processing by eliminating the need to handle data sequentially, allowing for parallel processing and significantly enhancing efficiency. The attention mechanism lies at the core of transformers, enabling the model to focus on the most relevant parts of the input sequence and improving accuracy and understanding of context. Furthermore, transformers excel at providing context, ensuring that the meaning of each word or token is accurately interpreted within its surroundings. Lastly, these models dramatically speed up the training process, making them faster and more efficient compared to traditional neural networks, thus redefining AI's capabilities across diverse applications.

Now that we know why transformers are important, let's discuss some applications. We have OpenAI's GPT, a groundbreaking model that leverages the power of transformers for natural language processing tasks. Additionally, Google has developed several transformer-based models, including Vision Transformer for image recognition, BERT (Bidirectional Encoder Representations from Transformers) for understanding the context of words in a sentence, and T5, which stands for Text-to-Text Transfer Transformer, for a wide range of text generation tasks. Microsoft has also contributed with DeBERTa (Decoding-enhanced BERT with disentangled attention), a model designed to improve contextual understanding and enhance NLP applications. These models demonstrate the versatility and impact of transformer architecture across various domains.

Now that we know the applications of transformers, how about checking their real-time products? Transformers have become an integral part of many real-world products that we use daily. Examples include Grammarly, which leverages transformers for advanced grammar and writing assistance; Google Search and its translation tools, powered by models like BERT and T5; and ChatGPT, OpenAI's conversational AI that relies on the Generative Pre-trained Transformer architecture. Additionally, Meta's deepfake detector uses transformer-based models for facial recognition tasks. These applications highlight how transformers have revolutionized technology, seamlessly integrating it into tools that enhance our everyday lives. In conclusion, transformers are changing the tech world by enabling smarter, faster, and more efficient AI systems. Whether it's generating text, translating languages, or enhancing search engines, these models are the cornerstone of modern AI.

So what are RNNs, right? Well, RNN basically stands for Recurrent Neural Network, and we usually use this in order to deal with sequential data. Sequential data can be something like a time series data or a textual data of any format. So why should one use RNN, right? Well, this is because there's a concept of internal memory here. RNN can remember important things about the input it has received, which allows them to be very precise in predicting what can be the next outcome. So this is the reason why they are performed or preferred on a sequential data algorithm. Okay. And some of the examples of sequential data can be something like time series, speech, text, financial data, audio, video, weather, and many more.

Although RNNs were the state-of-the-art algorithm for dealing with sequential data, they come up with their own drawbacks, and some of the popular drawbacks over here can be like, due to the complication or the complexity of the algorithm, the neural network is pretty slow to train. And as there are huge amounts of dimensions here, the training is very long and difficult to do. Okay. Apart from that, the most decisive feature for RNN or for the improvement in RNN is that of a vanishing gradient. What this vanishing gradient is, is that, you know, when we go deeper and deeper into our neural network, the previous data is lost. This is because of a concept called as vanishing gradient. And due to this, we cannot work on a large or a longer sequence of data. Okay.

To overcome this, we came up with some new or upgrades to the current recurrent neural networks or RNNs. Starting off with bidirectional recurrent neural network. You see, bidirectional recurrent neural network connects two hidden layers of opposite direction into the same output. With this form of generative deep learning, the output layer can get information from past future states simultaneously. So as you can see here, we have two layers over here, and as they are bidirectional, what happens is, when the algorithm feels that it is kind of losing its gradients or the previous data, it can go back and get the data.

from the past. So why do we need bidirectional recurrent neural network? Well, bidirectional recurrent neural network duplicates RNN processing chain so that the input process both forward and reverse time order, thus allowing bidirectional recurrent neural network to look into future context as well.

The next one is Long Short-Term Memory. Long Short-Term Memory, or also sometimes referred to as LSTM, is an artificial recurrent neural network architecture used in the field of deep learning. Unlike standard feed-forward neural networks, LSTM has feedback connections. It can not only process a single data point but also the entire sequence of data. So, as you can see here, from what I'm trying to say is, with LSTM or Long Short-Term Memory, it has something like, you know, we can feed a longer sequence compared to what it was with bidirectional RNN or RNN.

So why is LSTM better than RNN? We can say that when we move from RNN to LSTM, we are introducing more and more control over the sequence of the data that we can provide. The LSTM gives us more controllability, and there are better results. All right.

So the next type of recurrent neural network is the Gated Recurrent Neural Network, or also referred to as GRUs. You see, GRU is a type of recurrent neural network that is, in certain cases, advantageous over Long Short-Term Memory. GRU makes use of less memory and also is faster than LSTM. But the thing is, LSTMs are more accurate while using longer data sets. I'm sure by now you might have got a hint about the trend that has led to the improvement, right? So the trend over here is, you know, the model should be capable of remembering and taking in on a longer input sequence.

The game-changer part for the sequential data was developed when we came up with something called as Transformers. And this paper was something which is based on a concept called as "Attention is Everything." All right. So let's take a look at this. The paper "Attention Is All You Need" introduces a novel architecture called as Transformers. Like LSTM, Transformers is an architecture for transforming one sequence into another, while helping other two parts, that is, encoders and decoders. But it differs from previously described sequence-to-sequence models because it does not work like GRUs. Okay. So it does not implement recurrent neural networks. Recurrent neural networks until now were one of the best ways to capture the timely dependence on a sequence. However, the team presenting this paper, that is "Attention Is All You Need," proved that an architecture with only attention mechanisms, not using RNN, can improve its results in translation tasks and other NLP tasks. One of the best examples for Transformers is Google's BERT.

So what exactly is this Transformer? Right? You see here, we have an encoder on the top and a decoder on the bottom. Both encoder and decoder are comprised of modules that can stick onto the top of each other multiple times. So what happens here is the inputs and outputs are first embedded into an N-dimensional space since we cannot use this directly. So we obviously have to encode our inputs, whatever we are providing here. One slight but important part of this model is the positional encoding of different words. Since we have no recurrent neural network that can remember how sequences are fed into the model, we need to somehow give every word or part of our sequence a relative position. Since a sequence depends on the order of the elements. Okay, these positions are added to the embedded representation of each word. All right. So this was the brief about Transformers.

So let us now move ahead and see some of the popular language models that are available in the market. All right. So let us now start off by understanding OpenAI's GPT-3. The successor to GPT and GPT-2 is GPT-3, and it is one of the most controversial pre-trained models by OpenAI. The large-scale Transformer-based language model has been trained on 175 billion parameters, which is 10 times more than any previous non-sparse language model. The model has been trained to achieve strong performance on many NLP datasets, including tasks like translation, answering questions, as well as several other tasks.

Then we have Google's BERT. BERT stands for Bidirectional Encoder Representations from Transformers. It is a pre-trained NLP model which was developed by Google in 2018. With this, anyone in the world can train either their own question-answering module with up to 30 minutes on a single Cloud TPU or a few hours using a single GPU. The company then released this, showcasing the performance of 11 NLP tasks, including the very competitive Stanford Question Answering Dataset. Unlike other language models, BERT has only been pre-trained on 250 million words of Wikipedia and 800 million words of BookCorpus and has been successfully used as a pre-trained model in deep neural networks. According to researchers, BERT has achieved 93% accuracy, which has surpassed any previous language models.

Next, we have ELMo. ELMo, also known as Embeddings from Language Models, is a deep contextualized word representation that models syntax and semantics of words as well as their logistic context. The model, developed by Allen Institute for AI, has been pre-trained on a huge text corpus and learned functions from bidirectional models, that is, BiLM. ELMo can easily be added to their existing models, which drastically improves the features of functions across vast NLP problems, including answering questions, textual sentiment, and sentiment analysis.

Now let's answer the fundamental question that is, what exactly is Generative AI? Generative AI refers to algorithms capable of creating new content, whether text, images, audio, or even videos. It's like having a creative AI assistant that can take a simple input and produce engaging outputs. For example, GPT and Llama can write essays or code, while image generation models like DALL-E and Stable Diffusion can visualize unique scenes from descriptions. But let's look at some of the popular tools driving this innovation. Well, some of the standard tools in Generative AI include GitHub Copilot, which assists developers with code suggestions, and ChatGPT for text-based interactions. Image generation tools like Stable Diffusion and Midjourney help creators bring visual concepts to life. Google's Gemini merges text and image capabilities, while Adobe Firefly extends AI's reach to creative suites. So if you want to know how to use these tools, then check out our Generative AI Examples video, link in the description.

So you might wonder where are these tools being applied. Now let's explore them. Generative AI is transforming multiple creative fields. Image generation tools power visual design. Music composition algorithms create original scores, and AI assists video editors in automating tasks. LLMs help generate and translate text, while code generation tools like GitHub Copilot boost developer efficiency. AI-generated voices are even being used in audiobooks and voice assistants.

So now let's take some of these tools and check. So this time we will use Pictory AI and Flicky AI. First, let's explore Pictory AI. So for that, let's go to its site and check its functions. So we are at the Pictory AI site. And on the left side, we have the Home, Projects, and Brand Kits. And on the main screen, we have different features Pictory AI provides. So let's choose Text to Video. Here, let's write some names and descriptions and press Generate. Pictory AI is a tool designed for video creators that helps transform long-form content, such as articles or blog posts, into short, engaging videos. It uses AI to automatically extract key highlights and create professional-looking videos with minimal effort. Due to its simplicity and time-saving capabilities, Pictory AI is popular for social media content creation and marketing.

Now let's see our next tool, which is Flicky AI. So now we are at the flicky.ai site. Here we have different features like Videos, where you can create videos from all of these blogs, prompts, etc. You can also create Audios from these features, and then we also have a Design feature. And on the left-hand side, you can see options like Files, Templates, Brand Kits, Voice Clones, etc. So now let's take an idea and convert it into a video. Now let's write our topic and generate. Flicky AI is a content creation tool that turns text into videos using AI-generated voices and visuals. It helps users create professional videos quickly by pairing written content with stock images, animations, and voiceovers. Flicky is ideal for marketers, content creators, and educators looking to create engaging video content efficiently.

Now that we have seen the applications, so let's step back and look at the journey that brought us here. So basically, our journey starts in 1947 with Alan Turing's concept of intelligent machines. By 1961, Joseph Weizenbaum introduced ELIZA, the first chatbot. The 1980s saw the birth of recurrent neural networks, while 1997 brought Long Short-Term Memory networks to tackle sequential data, and then GANs emerged in 2014, transforming creative tasks. Fast forward to 2017, when Transformers like GPT entered the scene. By 2023, GPT-3.5 and Google's PaLM marked significant milestones, and by 2025, we are on the brink of AI breakthroughs in chemistry and genome editing.

So what exactly are these LLMs and why are they so powerful? An LLM, or Large Language Model, analyzes and understands natural language using machine learning. Examples include OpenAI's GPT, Google's PaLM, and Meta's Llama. These models drive applications such as chatbots, language translation, and more by learning from extensive data to predict and generate text sequences.

But before this, there was a very famous term called Language Model. A language model is a machine learning model that uses probability, statistics, and mathematics to predict the next sequence of words. Suppose you have a sentence like, "I have a boy who is my dash." Here, if we ask a language model to predict the next word, it considers the context provided by the words before the blank. Based on common usage patterns from its training data, it may predict words like "boyfriend," "brother," or "friend," which fit naturally. However, it's less likely to predict "colleague" or "sibling" as those words may not commonly follow these types of phrases. So this process shows how language models predict text by calculating probabilities for each possible word based on their likelihood in context.

So when a language model is trained on massive amounts of diverse text, it gains a wider vocabulary and more understanding of language, enabling it to make more accurate predictions. For example, if we give it a phrase like, "You are a dash to me," a model trained on extensive data might suggest various fitting words, for example, "friend," "inspiration," or anything else, so based on the sentiment or context it has learned from the data.

Now, here, reinforcement learning is used to improve the model's responses over time. By giving feedback, be it positive or negative, on the responses, we help the model learn which types of responses are preferred in specific contexts. For example, if the model frequently misinterprets the tone or intent, the reinforcement learning helps adjust its predictions to be more contextually appropriate and aligned with the intended meaning.

But what do these models look like under the hood? Well, LLMs are built on neural networks composed of input, hidden, and output layers. The hidden layers process information to learn complex patterns, and more layers mean the model can capture deeper insights. This structure allows LLMs to perform tasks from generating text to complex code completions.

Now, how do these layers interact and function in real-time? Now, an LLM is based on the Transformer, and a Transformer uses deep learning to process any information coming to it. Now let me tell you a story of three friends. Imagine we have three characters. First is our friend. The next character is Minion Bob. And the third character is Gru. So our friend asks Bob, "What's the price of the jet? It must be $50,000." Minion Bob isn't sure. So he goes to Gru and asks, "Is the jet $50,000?" Gru replies, "No, it's $70,000." In this back and forth, Minion Bob is like the neural network layer trying to make an accurate guess. So each time he goes back to Gru, like receiving more data or feedback, he gets corrected if his guess is wrong, leading him to refine his response.

Now, after the first check, Minion Bob returns to our friend saying, "I guess it's more than $60,000." Our friend assumes it might be around $65,000 and sends Bob back to Gru to verify. Again, Gru corrects him, "No, it's actually $70,000." So this process repeats with Bob adjusting his guess each time. Eventually, he learns that the correct answer is $70,000 and updates his knowledge. So just like Minion Bob, neural networks make initial guesses based on available information, with each feedback loop, like Bob going back to Gru, the model's hidden layers adjust the parameters to refine its guesses, ultimately arriving at the most accurate prediction possible. So after getting corrected multiple times, Minion Bob's guesses improve until he knows the price is $70,000. Similarly, in a neural network, gradually learning the correct answer through training. So once the network learns, it can give accurate answers in future cases without checking every time.

Now let us move on to understand how LLMs work. LLMs begin with the collection of datasets, then tokenize text and break it into manageable pieces. Using a Transformer architecture, they process the data sequence all at once, leveraging vast training data. LLMs contain millions of learned parameters that predict text tokens and generate coherent outputs. Models often undergo pre-training for general knowledge and fine-tuning for specific tasks.

So now let's see some practical uses of LLMs. LLMs power content generation, creating anything from articles to code. They excel in language translation, enhanced search engines, personalized recommendations, code development assistance, and sentiment analysis, which also owe much to LLMs' predictive capabilities.

So guys, are you ready to use all that knowledge in coding and witness how these LLMs come together to drive innovation? Whether through developing applications, analyzing data, or building smart assistants, the gear of technology keeps turning to unlock AI's full potential.

So now let us look at our problem statement. So one of the difficulties in the healthcare industry is effectively evaluating medical pictures, such as MRIs, CT scans, and X-rays, in order to identify anomalies and illnesses. This procedure takes a lot of time and calls for specialized understanding. Automated methods must be developed to help medical personnel recognize possible health problems in medical imaging. In order to provide better patient care, a system that integrates cutting-edge machine learning models with image analysis can greatly help in the early detection of diseases, including cancer, infections, and other illnesses.

So, the method uses generative AI to evaluate medical photos and generate a thorough diagnosis report based on the findings. This technology allows users to upload medical images, which the AI model then processes. Now let us build our project on a medical image analysis application using Streamlit, Python, and an LLM of Google Gemini AI. So this app helps healthcare professionals analyze medical images, such as X-rays, MRIs, and CT scans, to detect anomalies and diseases.

First, let's import the necessary libraries. So first, import streamlit as st. So if this is not working or showing an error, then open the terminal and write `pip install streamlit`. And from pathlib import Path. Next, import google.generativeai as genai. So we are importing streamlit for the app interface and Path from pathlib for handling file paths and Google generativeai, which allows us to interact with the Gemini AI model.

Next, we will configure Google's Gemini API by setting up our API key. So this will allow us to connect to the AI model and generate insights from medical images. So before proceeding, let's get our API key, and we will go to Google to generate an API key. So on your left, there is an API key option, and after clicking, you will get the "Create API" option. So just select your model and create your API key. So as you can see the screen, just copy this API key and go back to the terminal.

So now let's configure our model. So just type `genai.configure` and inside the bracket, give `api_key=` and over here, paste the key. Now, we set up the system prompt, which defines the role of the AI model. So the prompt specifies that our AI is a medical image analysis system capable of detecting diseases like cancer, cardiovascular issues, neurological conditions, and more. So guys, I have already researched the prompts and written here. So basically, the system prompt should be inside the triple quotes. So this prompt guides the model to analyze medical images for conditions such as cancer, fractures, infections, and more, making it a valuable tool for healthcare professionals.

Now let's configure the model settings for generating responses. We define parameters like temperature and top_p to control the creativity of the model's output. First, type `generation_config =` and inside the double quotes, we will give `temperature` which is 1, then `top_p` which is 0.95, next `top_k` 40, then `max_output_tokens` which is 8192, next `response_mime_type` which is of `text/plain`. So over here, the temperature 1 controls randomness; a value of 1 gives balanced output diversity. Next, top_p 0.95 uses nucleus sampling, selecting tokens from the top 25% cumulative probability distribution for diverse responses. Next, the top_k 40 limits token selection to the top 40 tokens based on probability, narrowing possible outputs to high-probability tokens. Next, max_output_tokens. This setting allows for longer responses by limiting the maximum length of the generated text to 8,192 tokens. And then we have response_mime_type, which specifies the format of the output as plain text. So for more information, read the Google Gemini documentation.

Next, we will also configure safety settings to ensure that the model doesn't generate harmful content. So, for example, we block categories like harassment, hate speech, and sexually explicit content. Here we are using two things: first, categories, and then the threshold. Then copy this four times, like harassment, hate speech, and sexually explicit content.

Now let's set up the layout for our Streamlit application. So for that, we will configure the title and the layout of the page and even add a logo to make the interface more user-friendly. So first, type `st.set_page_config` and inside the bracket, let's give `page_title=` and inside the double quotes, we will give "Diagnostic Analytics", comma `page_icon=` equal to "robot". Now let us type `column1, column2, column3 = st.columns(3)`. Next, with `column2`, I'll be using `edurea` and `medical images`. So this will show you how to set up images using Streamlit. Now type `st.image` and give a bracket, and inside the double quotes, let's type `edurea.png` and give a comma and give `width=200`. Now let us copy and paste it for medical. So let's type `medical.png`. Here we are using Streamlit's columns to center the logo and title, and this makes the app look professional and visually appealing.

Next, let's allow the user to upload medical images for analysis. So, we use Streamlit's file uploader widget to accept image files in PNG, JPG, or JPEG formats. For that, let's type `uploaded_file = st.file_uploader` and inside the bracket, inside the double quotes, let's type "Please upload the medical images for analysis.", comma `type=` so basically the image type is equal to and inside the bracket, inside the double quotes, let's give `["png", "jpg", "jpeg"]`. Next, let us type `submit_button = st.button` and inside the bracket, let's give "Generate Image Analysis". So here, when the user uploads a file and clicks the "Generate Image Analysis" button, the model processes the image and prepares it for analysis. So once the user submits the image, we send it to the AI model for analysis, and then the model generates a response based on the prompt and image, which we then display in the app.

So here, as you can see the screen, we have another function. So the `if submit_button:` which runs the code when the submit button is pressed. Next, `image_data = uploaded_file.getvalue()`. This actually gets the raw image data from the uploaded file. And next, we have the `image_parts` where it creates a list with the image data in a structured format. Then we have the `prompt_parts`. So this combines the image data and a text prompt for the model. So this part of the code actually sends the image and text prompt to the model to generate a response. And then we have the `st.write` which displays the model's responses in the app. So here we use the image data and system prompts to generate content with the Gemini AI model. The result is displayed as a detailed report with insights about the medical image.

Now it's time to test the code. So open the terminal and type `streamlit run main.py`. So once you enter it will redirect you to our model interface. And there you go. So the model is ready. So here's a live demo of the app. We will upload a sample image, and the app will analyze it and provide a detailed diagnosis based on the AI models inside. So this is how we use Streamlit and Google's Gemini AI model to create a medical image analysis app. So this app can help medical practitioners by offering precise and thorough analysis of medical photos.

Now it is the time for testing. So let's take one image of any disease and test it. So upload the image from your computer. Then we will select an image and press the generate button. So as you can see, it's running. So it generates a fabulous response and can help doctors in assisting their patients, saving time and money. So this is how we built a real-time medical diagnostic helper using Streamlit, Python, and Google Gemini AI.

Have you seen how tools like ChatGPT with Vision can look at an image you upload and describe it? Or how DALL-E and Midjourney can generate stunning images from just a text prompt? And now some AI models can even do both at the same time. They can see, read, listen, and even create, all in one go. So how is that possible? Well, that's because of something called multimodal AI. AI that doesn't just work with one type of data, like only text or only images, but can understand and combine multiple types of information together, just like we humans do. So, in this video, we are going to break down what multimodal AI really means, and how multimodal AI works, and explore some amazing real-world examples that you're probably already using without even realizing it.

So first, let's break down the word multimodal. So "multi" means many, and "modal" refers to the modes of information, like text, images, sound, or video. So multimodal AI is an AI that can understand and work with multiple types of data at the same time. For example, a single AI model that can read text, look at images, listen to audio, watch videos, and combine all of this to give a better answer. It sounds a bit like how humans process information, right?

So, why do we need multimodal AI? So think about how we interact with the world. If you're watching a movie, you're seeing visuals, listening to dialogue, and understanding the story together. Or when you're explaining a recipe to someone, you might show pictures, describe steps, and maybe even play a video. So humans naturally combine different senses to understand things. So, old AI models were single-modal. They could only process one type of data. Like a text model could only read and write, and a vision model could only look at images. But real-world problems are not just text or just images. They are mixed. So multimodal AI bridges this gap, and it lets AI connect the dots between text, visuals, audio, and more.

So how does multimodal AI work? In simple terms, it works like this. It takes different types of input. For example, it could take a photo and a text question about that photo. Then it converts them into a common language inside the AI model. So think of it like translating text, images, and audio into one shared understanding. Next, it reasons over all the data together. Then it gives you a smart answer that considers all the inputs. So, for example, you show AI a picture of a dog and ask, "What breed is this?" So it looks at the image, understands the features, and responds, "That looks like a Golden Retriever." So it's combining vision plus language to answer.

Now let us go through a working diagram of a full multimodal pipeline. So as you can see the screen, first it takes different inputs. It could be a text, image, or even a video. Then it has encoders for each modality. Later, these inputs will be translated into a common AI language. Then a multimodal Transformer uses cross-attention to connect relationships across text, images, and audio. And finally, the model generates a response. So let me take another example to explain this diagram. So as you can see, we have different inputs. So the model can take text, images, audio, or even video as input. Next is the encoders for each modality. That means a text encoder converts words into vectors, and an image encoder converts pixels into vectors, and then an audio encoder converts sound waves into vectors. Next is the shared embedding space, where all the different inputs are translated into a common AI language, which is a vector space where similar meanings are clustered together. For example, the word "car" and a picture of a car are mapped close together. Next is the fusion plus reasoning layer, where a multimodal Transformer uses cross-attention to connect relationships across text, images, and audio. For example, it links the word "red" to the red region of the car image. Next is the output generation. So finally, the model generates a response, which could be text, a caption, an image like DALL-E, or even sound.

All right, I hope this is clear now. So now let's look at some real-world examples that make it easier to understand. So first, we have ChatGPT with Vision. So if you upload an image to ChatGPT and ask, "What's in this picture?" then it can describe the objects, text, or even analyze data like a chart. So that's multimodal AI. It's using both image understanding and text generation together. The next example is Google Lens. So when you point your camera at something, Google Lens can recognize the object, read the text in the image, and translate it into another language. Again, it's a vision plus language plus translation, all in one model. The next example could be self-driving cars. So autonomous cars like Teslas use multimodal AI because they have to see the road through cameras, read traffic signals, hear alerts, and also process maps and text instructions. So they combine all these modes to make driving decisions. Next is healthcare AI. So doctors now use AI that can look at medical images like X-rays and also read patient reports, combining the information to help diagnose diseases more accurately.

But why is multimodal AI a game-changer? Multimodal AI is powerful because it's closer to human intelligence. We don't rely on one sense; we combine many. And it makes AI more flexible because one model can handle text, images, audio, and more. It can solve more complex problems like explaining what's happening in a video or understanding a full conversation with context.

All right. Now, for those of you who want a bit more technical depth, here's a quick peek behind the scenes. So, as I discussed earlier, multimodal AI uses Transformer-based models, the same type of models behind GPT. So the text, images, and audio are all converted into common representations, like a shared language of numbers called embeddings. For example, a picture of a dog and the word "dog" are both mapped into a similar space. So the AI knows they mean the same thing. Then the model can reason across all modalities together and generate an output. A great example is CLIP from OpenAI, which connects images and text. Another is Google Gemini, designed from the ground up as a truly multimodal model.

So what is the biggest challenge? So the different types of data have different formats and complexity. Combining them efficiently without losing meaning is still an ongoing research area. So it's not just magic. So it's smart design that lets the AI translate everything into one common understanding.

Let's now look at some of the most important multimodal models, how they work, and where they are used. So here are the key multimodal models. So first on the list, we have CLIP, which is CLIP, which stands for Contrastive Language-Image Pre-training from OpenAI. So let's see how it works. So it has two encoders: a text encoder and an image encoder. So both encoders map inputs into the same embedding space. So during training, it learns this caption matches this image, and this caption does not match that image. So it uses contrastive learning. It pushes correct pairs closer and incorrect pairs further apart. So here is the working diagram. So it takes the input, be it image or text, and then it has encoders. So an image is a vision encoder, and for text, it's a text encoder. Then it is shared to an embedding space, and finally, it generates the output. So let's have a look at the use cases. So it is used in DALL-E and Stable Diffusion to align text prompts with images. Next, it is used in zero-shot classification, where you give it a photo of a dog versus a photo of a cat, and it recognizes which one matches the image without retraining. And then it is used in search, where it finds images similar to this caption.

Next, moving on to the second model, which is BLIP-2. So it stands for Bootstrapping Language-Image Pre-training. So let us see how it works. So first, it connects a frozen vision encoder, for example, CLIP or ViT, with a frozen Large Language Model, which is an NLM. A Query Transformer acts as a bridge, where it converts visual features into a language-friendly representation. So here is the working diagram. The AI first looks at the image and turns it into features like objects, color, and shapes. Then a small bridge model called Q-Former takes those visual features and converts them into a format the language model can understand. Next, the Large Language Model then reasons about the image features, just like it reasons about text. And finally, it generates a text answer or a caption describing the image. So, visual encoder sees, Q-Former translates, and the LLM explains. So let's have a look at the use cases. So first, it is used in visual question answering, for example, "What's in this picture?" Next, in image captioning, where it can give, like, "A man riding a horse on a beach." Next, in chatbots with vision, for example, where you upload an image and ask questions.

Okay. The next model on our list is Flamingo from DeepMind. So let's have a look at its working. So here's how it works. So first, it's a few-shot multimodal model. It doesn't need huge fine-tuning for a new task. And then it uses gated cross-attention layers to integrate image plus text inside a frozen LLM. And it can reason across multiple images and a long text sequence. So it looks at the image, reads your question, connects both through cross-attention, and then explains it. So let's have a look at the use cases. So it is used in multimodal chatbots, like, "Look at these five images and now answer this question." Next, it is used in educational AI, where it reads diagrams plus answers questions. Next, in document understanding, where it reads text plus images in a PDF.

And the next multimodal on a list is PaLM-E from Google. So here's how it works. So as you can see, this is the working diagram. So first, the AI gets both visual input, like a photo or a live camera feed, and the text instructions, like "Pick up the red apple on the table." Next, the vision transformer understands what's in the image, like objects, colors, and position, and the PaLM language model understands the instruction and reasons about what needs to be done. So it's combining both. The AI creates a step-by-step action plan for the robot, like "Move forward, grab the red apple, and place it in the basket." So, here are the use cases. It is used in robotics, like "Pick up the red apple on the table." Next, it is used in real-world reasoning for embodied AI. Then, it is also used in visual navigation tasks.

And the next multimodal is Google Gemini. So here's how it works. It's natively multimodal, trained from scratch on text, images, audio, and video. So unlike CLIP, which aligns two encoders, Gemini has a single model handling all modalities, and it uses joint training with cross-attention. So this is the working diagram. So let me explain this. The AI takes in all types of inputs at once, such as written text, pictures, sound, and even video. Then, instead of using separate models for each type, it uses one powerful Transformer model that can understand and combine all these inputs together. And from that combined understanding, it can give any kind of output: a text answer, a generated image, or even an audio response. So basically, it understands everything together and responds in any form you need. So let us have a look at its use cases. It is used in complex queries, such as "Summarize this video and create a chart." It is used in advanced digital assistants and also in future AR/VR multimodal applications.

The next model is GPT-4o from OpenAI. It's an optimized multimodal model. It accepts text, images, audio in real-time, and it uses fused embeddings and parallel processing for speed, and it works as a true interactive assistant. So here are its use cases. It is used in conversational AI with vision plus audio, and in real-time assistance where you upload an image and get an explanation instantly, and also in accessibility tools, for example, "Describe surroundings for visually impaired users."

So these models represent different approaches to multimodality. Some align separate encoders like CLIP. Some bridge vision plus LLMs like BLIP-2, and some are natively multimodal like Gemini and GPT-4o.

So now let us see how are multimodal models trained. So training multimodal models is much more complex than training single-modal models. So first is the data set alignment. So you need paired datasets, such as images plus captions, videos plus transcripts, and audio plus text. So the challenge is the text and images don't always align perfectly. Next is the contrastive learning. So train the model to pull matching pairs closer and push non-matching pairs apart. For example, an image of a cat plus the caption "a cat" is a matching pair, whereas an image of a cat plus the caption "a dog" is not a matching pair. Next is the masked modeling. Mask parts of the input, such as image patches or text tokens, as the model predicts missing information. Then it forces the model to reason across modalities. For example, mask the object in a caption, "A dash is sitting on the table," plus providing an image. Next is the fusion and cross-attention training, where models like Flamingo or Gemini train cross-attention layers to integrate modalities. It requires huge compute clusters. Next is scaling loss. Like LLMs, multimodal models get better with size and data diversity. Gemini and GPT-4o are trained on massive multimodal corpora.

So here are the training requirements. You need to have high-quality paired datasets, billions of parameters, and TPUs, GPUs for weeks or months, and advanced optimizations such as mixed precision or shared training.

So why is true multimodal AI still hard? It's because of data mismatch. Text is sequential, images are spatial, and audio is temporal. So aligning them perfectly is difficult. Next is limited high-quality data. So billions of image-text pairs exist, but have noise and bias. Next, bias and fairness. Models learn cultural and social biases from multimodal data. For example, stereotypes in images and captions. The next challenge is compute cost. So training needs huge GPU clusters, for example, hundreds of A100 GPUs, and fine-tuning multimodal models is even more expensive than text-only. And the final challenge is the evaluation difficulty. So how do you measure reasoning across modalities? So there's no single easy benchmark. So while multimodal AI is powerful, it's also data-hungry, compute-heavy, and still evolving.

So in simple words, multimodal AI can process and combine multiple types of data, such as text, images, audio, and video. It's already in use, such as GPT Vision, Google Lens, self-driving cars, and healthcare AI. It's a big step towards AI that can understand the world like humans do. So what do you think? Will multimodal AI make AI more humanlike? So drop your thoughts in the comments.

LLMs like GPT-4 and Gemini 2.0 are massive models trained on huge datasets, capable of generating highly sophisticated and nuanced responses. On the other hand, SLMs like DistilBERT or TinyGPT are smaller, more efficient models designed for faster and more lightweight tasks. So understanding the differences between them is crucial for selecting the right model for your needs. Now let's dive right in with our first question.

What exactly are LLMs and SLMs? LLMs, which are Large Language Models, are powerful AI systems trained on vast datasets, offering deep contextual understanding and sophisticated responses. So models like GPT-4 and Gemini 2.0 are examples. Whereas SLMs, like DistilBERT or TinyGPT, are streamlined for speed and efficiency, excelling in lightweight tasks. So both serve distinct purposes, balancing quality, cost, and performance.

All right. Now that we have got a good idea of what LLMs and SLMs are, let's talk about why this comparison is so important. As AI adoption grows across industries, the choice between LLMs and SLMs becomes more important. LLMs offer deep contextual understanding and complex outputs, while SLMs provide efficiency and speed. So choosing the wrong model can lead to excessive cost, slow performance, or subpar results. And by understanding the strengths and trade-offs, you can make more informed decisions and optimize your AI solution.

So now let's dive into the core differences between LLMs and SLMs and see what sets them apart. So first, let us compare in terms of model size and complexity. So when it comes to model size and complexity, LLMs often have billions of parameters and require vast computational resources to train and run. Their large size enables them to generate high-quality, contextual responses. And on the other hand, SLMs are designed with fewer parameters, often in millions, making them lighter and faster. They prioritize efficiency over complexity, which makes them ideal for simpler tasks.

Next, let us compare in terms of performance and output quality. So, when it comes to performance and output quality, LLMs are known for their exceptional ability to handle complex conversations, creative writing, and deep analysis. Their vast training data ensures diverse and sophisticated responses. On the other hand, while SLMs are efficient, they may sometimes struggle with nuanced or open-ended queries. However, they excel in straightforward, well-defined tasks.

Next, let's compare them with speed and latency. When it comes to speed and latency, LLMs can experience longer response times and higher latency due to their large size, especially when processing extensive input data. Whereas SLMs are designed for speed, offering quicker responses and making them well-suited for real-time applications where low latency is crucial.

Next, in terms of cost and resource efficiency. So when it comes to cost and resource efficiency, LLMs require significant hardware investments, such as powerful GPUs and extensive cloud resources, which lead to higher operational costs. Whereas SLMs, with their smaller footprints, are more affordable to deploy and maintain, making them accessible even with limited computational resources.

Now let us explore the real-world use cases of LLMs and SLMs. LLMs are ideal for creative content generation, customer service chatbots with advanced capabilities, deep data analysis, and long-form conversations. On the other hand, SLMs are perfect for lightweight virtual assistance, real-time customer support, simple automation, and tasks that require quick turnaround times.

Now, let us see its advantages and disadvantages of using LLMs and SLMs. The key advantages of LLMs include their superior understanding of complex language, the ability to generate high-quality, nuanced responses, and better generation across a wide range of diverse tasks. The main drawbacks of LLMs are their high computational and cost demands, along with slower response times due to their large size and complexity.

Now let us have a look at the advantages of SLMs. SLMs offer several advantages, including their speed and efficiency, lower operational costs, and easier deployment, even on limited resources. The primary disadvantages of SLMs are their limited contextual understanding and their tendency to struggle with complex, open-ended queries.

Now that we have explored the strengths and limitations, so let's take a look at what the future holds for LLMs and SLMs in AI development. So both LLMs and SLMs will play a vital role in the future of AI. We can expect ongoing improvements in efficiency, quality, and adaptability. Hybrid approaches that combine the strengths of both models could become more common, offering balanced performance and scalability.

So the conclusion we get is that the choice between LLM and SLM depends on your specific needs. So if you prioritize depth, nuance, and high-quality output, LLMs are the best. So if speed, efficiency, and cost are more important, SLMs are the way to go. So by understanding their strengths and limitations, you can select the right model and unlock AI's full potential for your projects.

Now let us understand what Langchain is and why it is a valuable tool for building AI applications. You must be aware of popular applications such as GPT and Gemini. These applications utilize APIs, and GPT uses OpenAI's API, while Gemini operates through the Gemini API to process prompts. They leverage models like GPT-3.5, GPT-4, PaLM, and Gemini 1. Additionally, there are other advanced models such as Llama, Gemini, Cohere, Cloud version 1, Falcon, PaLM, GPT-4, and GPT-3.5. Langchain is a framework designed to help developers build flexible and powerful AI-driven applications by integrating and utilizing these diverse models effectively.

But why exactly do we need Langchain? You must be thinking, if Langchain is this important, then why do we need Langchain? So let's break down this question using some real-world examples. So imagine simply asking an LLM a prompt and getting an answer. That's easy. But what happens when the complexity increases? For example, let's say you're working with data from SQL databases, CSV files, PDFs, or Google Analytics, and you need the model to write code, perform searches, or send emails. Handling such intricate workflows manually can get overwhelming. This is where Langchain steps in. It simplifies the process by offering components like document loaders, text splitters, vector databases, prompt templates, and tools. So this helps you assemble tasks such as document summarization, question-and-answer systems, or even advanced workflows like Google searches or customer support automation.

Let's visualize this process with a diagram. Here's how it works. So first, you load a document like a CSV file using a document loader. Then use a text splitter to divide it into smaller chunks, and then store those chunks into a vector database, and add a prompt template to guide the model. And finally, use an LLM like GPT-4 or Llama to perform tasks like searching the web or automating workflows. And Langchain also offers chains that will help you assemble components to achieve single tasks, such as summarization, and an agenda to figure out what each component must do, like password, customer services, etc.

Now that we understand Langchain's core components, now let's explore how it streamlines the LLM application lifecycle. So it typically involves three key stages. First is the development, where you build and test your application. Then productionization, where the system is fine-tuned for real-world use. And finally, deployment, where the final product is launched for users. So Langchain simplifies this lifecycle, allowing you to focus on building without worrying about the underlying complexity.

Now let's take a step back and understand the role of APIs in powering these LLM applications and how Langchain effectively integrates them. In all these applications and models, one thing is common: that is, they use APIs. So now let's discuss APIs. APIs act as intermediaries that enable different systems to communicate with each other. For example, they allow apps like Swiggy or Blinkit to display your delivery driver's location in real-time.

So now let's look at the steps to explain APIs and API keys. So apps like Zomato, Swiggy, and Blinkit use APIs to show the location of your delivery driver. So these apps don't communicate directly with Google Maps but follow a layered process involving servers and security mechanisms. First, the app sends a request to the Google Maps API. Then the API forwards the request to Google servers. Then the servers validate the request with the system. So once approved, the response follows back through the servers, APIs, and finally to the app. So previously, apps like Swiggy allowed login using phone numbers. Now they use APIs for login via platforms like Google or Facebook. So this demonstrates the versatility of APIs in enabling seamless user interactions.

To prevent misuse, APIs require API keys, which are unique identifiers for secure access. So these keys authenticate requests and ensure that only authorized users can interact with the APIs. Next, security systems closely monitor API usage to detect and prevent misuse. This ensures that APIs remain safe and functional for their intended purpose. And these steps simplify the explanation of how APIs and API keys work in real-world applications. So this is how Langchain leverages APIs to connect your LLM applications with external tools, making them versatile and secure.

Now that we understand the role of APIs, so let's explore some real-world applications of Langchain. So what can you build with Langchain? Here are a few applications. First application we have is customer support. So customer support for your shopping websites to interact with customers. Next, conversational chatbots for helping you study, content generation tools for blogs or social media. We also have question-answering systems for knowledge bases, and then document summarizers for legal or academic content. Langchain simplifies AI development by integrating LLMs with various data sources and tools. Its applications are vast, from chatbots to document summarization.

So, let's examine a practical example to see Langchain in action. All right. In today's data-driven world, understanding and effectively using SQL queries is crucial for managing and analyzing large datasets. However, beginners and even experienced users often need help with complex SQL queries, their syntax, and how they work. This creates

A barrier to efficiently interacting with databases and limits their potential to solve real-world problems. To address this challenge, we propose a SQL query fetcher application that leverages the Gemini AI, Python, and Streamlit to simplify SQL learning and usage. The application allows users to input or select a query, generates the SQL syntax, and provides a detailed explanation of its components and functionality. This tool bridges the gap between technical understanding and real-world database operations, empowering users with an initiative and interactive SQL learning experience.

Let's jump right into the code. So the first step is setting up your dependencies. Here we import Streamlit for the user interface and then Google Generative AI for using Gemini. So first, import Streamlit as st and next, import Google dot generative AI as genai.

So to get this API, you have to go to the Google Gemini API key and here, click on "Get a Gemini API key" in Google AI Studio. And then, once you scroll, there is a button on the left called "Create API". Now, click on it and select your model here and let's copy it. And now, let's go back to our VS Code editor and paste it here. So to paste, let's type `google_api_key` and inside the double quotes, let's paste it. And now, let's type `genai.configure()` and inside the bracket, let's keep it as `api_key=google_api_key`. Now, let's type `model = genai.GenerativeModel('gemini-pro')`.

So, we use the Google Gemini API to generate SQL queries dynamically. So, make sure to configure your API keys securely.

Now, let's display the Streamlit layout code. Now, let's set up the app's user interface. So, we use Streamlit to create an interactive page where users can input plain English queries and get SQL code in return. So, we write `st.set_page_config(page_title='Edureka SQL Query Generator', page_icon=':robot:')`.

Now, let's put some images. So, I'm using Edureka image and SQL logo. And also, to make them center, we will type it as `col1, col2, col3 = st.columns([1, 2, 1])`. Next, let us type `with col2:` and let's type `st.image(image_address, width=200)`. Now, let us add another image. So, let's copy the same and give the other image address.

Our layout includes a title, logo, and text input box to keep the interface simple and intuitive.

So, here's where the magic happens. So, when a user clicks the "Generate SQL Query" button, we format their input into a prompt for the Gemini model to generate SQL code. So, let's create a template by writing `template = """Create a SQL query snippet using the below text. Text: {text_input} SQL Query: """`. Next, let us also give `text_input = st.text_area("Enter your query here in plain English")`. Now, let's type the response. So, type `response = model.generate_content(template.format(text_input=text_input))` and let's keep it as `sql_query = response.text.strip()`.

So, the AI generates the SQL query, and we clean up the output for display. So, once the SQL query is ready, we take it a step further by generating a sample expected output and a clear explanation of the query.

Now, let's type the logic for showing explanation and output. So, let's type `st.markdown("""<h1>SQL Query Generator</h1><h3>I can generate SQL queries for you.</h3><h4>With explanation as well.</h4><p>This tool allows you to generate SQL queries based on your data.</p>""")`. Now, to make the markdown visible, let us type `unsafe_allow_html=True`.

Now, let's write `text_input = st.text_area("Enter your query here in plain English")`. Now, let us give a submit button. So, for that, let us type `submit_button = st.button("Generate SQL Query")`.

Now, if `submit_button`:

`with st.spinner("Generating SQL query..."):`

`template = """Create a SQL query snippet using the below text. Text: {text_input} SQL Query: """`

`response = model.generate_content(template.format(text_input=text_input))`

`sql_query = response.text.strip()`

`st.success("SQL query generated successfully!")`

`st.subheader("Here is your query below:")`

`st.code(sql_query, language="sql")`

`st.success("Expected output of this query will be:")`

`st.markdown(f"```\n{output}\n```")` # Assuming 'output' is defined elsewhere

`st.success("Explanation of this SQL query:")`

`st.markdown(f"```\n{explanation}\n```")` # Assuming 'explanation' is defined elsewhere

Over here, this shows a green success message indicating the SQL query was generated successfully. And next, the "Show SQL Query" displays the SQL query as a formatted code block, highlighting it as SQL. Next is the "Display Expected Output". This provides a success message for the query's expected output, followed by `st.markdown(f"```\n{output}\n```")` which displays the expected output in markdown format. So now, this line of code introduces an explanation, and `st.markdown(f"```\n{explanation}\n```")` displays that it is in markdown format for clarity. So, this makes the tool valuable for both learning and debugging SQL.

Now, let's see it in action. So, open the terminal and let us type `streamlit run your_file_name.py`. Now, as you can see on the screen, your SQL query generator is ready to go. Now, let's test it. So, for that, here I will input a prompt asking for a query, which is "Give me the query for create table". Now, let's click on "Generate SQL Query" and as you can see, it's running. So, let's wait for it to generate.

So, as you can see on the screen, the app generates a SQL query, expected output, and even a plain English explanation in seconds. So, how cool is that, right? And that's it. Our SQL query generator, powered by Langchain, Gemini API, and Streamlit, is complete. So, this project is perfect for simplifying SQL learning and enhancing productivity.

Before we talk about agents, let's quickly understand Langchain. Langchain is a framework designed to help you connect large language models, such as GPT, with external tools, APIs, memory, and custom logic. Normally, LLMs like ChatGPT can only generate responses based on the text you give them. But what if you wanted to search the web, run Python code, query a database, or use a calculator? That's where Langchain comes in. It acts as a bridge between the LLM and the tools it can use to interact with the real world.

And one of the most powerful features in Langchain is agents. So, what exactly is a Langchain agent? So, think of it like this: instead of you telling the AI exactly what to do, you just give it a goal, and the agent figures out how to get it done. An agent combines the power of reasoning, decision-making, and tools. It uses the LLM to understand the task, choose which tools it needs, call those tools in the right order, and then return the final result to the user. It's like giving your AI assistant a toolbox and letting it decide which tools to use based on the question you ask.

So, let's look at a real example to make it clear. Imagine this prompt: "Check the current stock price of Apple and calculate the average over the past 5 days." A regular chatbot can't do that. But a Langchain agent can use a web search API to find today's stock price and use a Python tool to calculate the average and then respond with the results, all automatically.

So, here's what's happening under the hood: The LLM receives your prompt and it decides it needs to search and calculate, and then it picks the right tools, maybe SER API for search and Python for math, and it performs each step in a sequence and gives you the final output. So, this is all done dynamically, meaning you don't hardcode each step. So, the agent figures it out using the language model's reasoning.

Now, let's look at the inner workings of a Langchain agent. So, when you create an agent in Langchain, you define three things: First, the LLM to use, like GPT-4 or Claude. Next, the tools available, like calculator, web search, database query, etc. Next, the agent type. So, Langchain supports types like Zero-shot Agent and Conversational Agent.

So, now that you understand how Langchain agents work, let's quickly talk about the two most popular types. So, first, we have the Zero-shot Agent. So, this is the most commonly used agent. It works by giving the language model a list of tools along with a description of what each tool does. Then, the model uses that information to figure out on the fly which tool to use and in what order. So, it's called "zero-shot" because the model doesn't get examples; it just reasons based on the tool descriptions. And it is best for tasks that don't need memory or a back-and-forth conversation, just like data lookups, calculations, or API calls.

And the next type is the Conversational Agent. This one is more advanced. It is designed for multi-turn conversations. That means the agent remembers previous steps and keeps track of what's already been done. So, it uses a chat history and a memory module to maintain context across multiple prompts. And it is best for chatbots, virtual assistants, or tools where the user asks follow-up questions or expects the AI to remember context.

So, in short, I can say that the Zero-shot Agent is fast, simple, and for one-shot tasks, whereas the Conversational Agent is context-aware, for back-and-forth dialogues. And there are other agents like Tool-Using Agents, Plan and Execute Agents, or Multi-Action Agents for more advanced workflows and perfect for future deep dives.

Now, to understand Langchain agents, let's quickly explore its core building blocks. So, first, we have the LLMs. It is the brain of the system. Langchain supports modules like OpenAI GPT-3.5 or GPT-4, also Anthropic Claude, and Hugging Face Cohere, etc. And the next component is prompts. These are the templates that guide the LLM's behavior. You can use static prompts or chat prompt templates for more dynamic and multi-turn interactions. Next, we have chains. It is a sequence of calls or logic. Next, tools. These are the external functions the LLM can call, such as Python calculator, web search API, or SQL query executor. Then, we have agents. Agents dynamically decide which tool to use and when, based on your input. So, agents are what make Langchain go from a chatbot into a multi-tool problem solver.

So, over here, the tools are just Python functions wrapped in a Langchain format. For example, this is the Python function. So, it's like giving your AI assistant a toolbox and letting it decide what to use based on your prompt.

Next, the agent then follows a process called ReAct, which stands for Reasoning plus Acting. So, here's what that looks like: So, first, the agent receives the prompt. Then, the LLM decides, "I need to search the web." So, the Langchain calls the search tool, and then the tool returns the result. Next, LLM reasons, "Now I need to do a calculation." So, the Langchain calls the calculator tool, and the agent returns the final answer.

So, here is the example of the ReAct loop: The thought is, "I need to find today's weather." The action that takes is, it uses a weather API. Next is the observation: "For example, it's 28°C in Bangalore." The thought is, "Now I can tell the user the temperature." And the final answer would be, "It's currently 28°C in Bangalore."

So, this entire flow is written and passed by the LLM itself using intermediate steps called scratchpads. So, the Langchain passes those steps and knows when to call a tool or stop.

Now, let's look at where Langchain agents are used in real-world projects. So, first, it is used in AI customer assistance. So, the agents can look up user info, reset passwords, and respond to queries automatically. So, users can ask things like, "What was my profit margin last quarter?" So, the agent pulls data from a database, does the math, and explains it. Next, the Langchain agents can be used in research tools. So, you can build a research bot that searches multiple sources, summarizes, and gives you an answer step by step. Then, in automated workflows like "send a message, create a task in Trello, and update the CRM," all with one prompt.

So, that's the power of Langchain agents. So, they allow your language models to take action, use tools, and solve real-world tasks step by step. So, let me know in the comments if you want a full coding tutorial on building your first Langchain agent.

RAG is a hybrid approach in artificial intelligence that combines retrieval systems with generative models to produce highly accurate, contextually relevant responses. It bridges the gap between factual accuracy and natural language generation. Now, let's understand it with the help of a diagram. So, it's a hybrid approach involving artificial intelligence that combines a retrieval system with a generative system to produce highly accurate responses.

Now that we know what RAG is, so let's explore why it is crucial for large language models and see a real-world example. So, RAG addresses several limitations of traditional LLMs. It mitigates hallucinations by grounding responses in factual retrieved data by dynamically accessing up-to-date information. RAG stays relevant in rapidly changing domains. It improves accuracy and relevance by fetching specific, relevant documents during inference. By outsourcing factual knowledge retrieval, RAG enables smaller, more efficient models, and it can adapt to domain-specific knowledge bases for specialized applications. Additionally, RAG provides explainability by showing the retrieved documents or data sources, increasing trust and transparency.

Now, let us see some of the use cases. So, without RAG, the sentence would be: "When was the last Mars rover launched?" So, this is just an incorrect response. So, with RAG, the sentence would be: "Dynamically retrieved from NASA's database, the Perseverance rover was launched on July 30, 2020."

Now that we have seen why RAG is important, so let's dive into how it works. Well, RAG operates in a three-step process. A user submits a query, which triggers the retrieval stage. Here, a retriever searches a database or knowledge base using tools like FAISS to fetch the most relevant information. The retrieved data is then fed into a generative model like GPT or T5, which processes it and generates a coherent, contextually grounded, natural language response.

Now, let's take an example here. The query is: "Who wrote 1984?" Retrieve would be fetching a document containing "George Orwell wrote 1984." Now, generative response would be: "The author of 1984 is George Orwell." This hybrid approach makes RAG ideal for real-world applications like chatbots and knowledge systems.

Now that we understand how RAG works, let's explore some of its real-world applications. RAG's versatile applications span various domains. In knowledge management, it can summarize large databases or documentation, aiding corporate teams. Legal and compliance tasks benefit from RAG's ability to answer queries based on case law and regulations. While in healthcare, it can support medical professionals by summarizing research papers and guidelines. Education and e-learning can leverage RAG for virtual tutoring, providing detailed explanations based on textbooks and research papers. Interactive virtual assistants like Alexa and Siri can utilize RAG to generate accurate and informative responses to user queries, such as news headlines or product recommendations. RAG's unique ability to combine retrieval and generation makes it essential for tasks demanding both factual accuracy and fluent natural language responses.

Now, let's compare Retrieval Augmented Generation with traditional AI models across three features. First, we have factual accuracy. RAG provides highly accurate responses by using real-time data, whereas traditional models may give less accurate answers and may give errors. Next is the context adaptability. So, here, RAG adapts quickly to new queries using live data, whereas traditional models offer fixed answers based only on pre-trained knowledge. Next, we have knowledge updates. RAG is easy to update; just change its data source, whereas traditional models need retraining, which takes time. Then, we have scalability. Whereas traditional models are limited by data size and training data. And then, we have use cases. RAG is great for tasks like legal advice or customer support, whereas traditional models work well for creative writing or casual queries. So, here I want to conclude that RAG is ideal for knowledge-based tasks needing accuracy and flexibility, while traditional models are better for creative users.

While RAG offers significant advantages, it's essential to acknowledge its limitations. So, let's discuss the challenges and future of RAG. So, the first challenge is the latency. RAG systems can suffer from latency issues, especially when dealing with large datasets or complex queries. Next is the data quality dependency. The quality of the generated responses heavily depends on the quality of the underlying data. The next challenge is complex integration. Integrating RAG systems with existing applications and infrastructure can be challenging due to the need for data synchronization, query optimization, and model management. And finally, scalability issues. As RAG systems become more complex and are deployed at scale, they can face scalability issues. This includes handling increased query loads, maintaining data freshness, and ensuring model performance.

Now, while RAG faces limitations, its potential is undeniable. So, now let's discuss RAG's future. The future of RAG holds immense potential. It will power dynamic, real-time applications like news summarization, financial analytics, and live sports commentary. RAG will be customized for specific domains like healthcare, law, and science through integrations with specialized knowledge bases. Advances in retrieval models and compression techniques will reduce latency to enhance efficiency. RAG will expand to handle multimodal data, enabling use cases like multimedia question answering. Additionally, RAG will facilitate personalized AI assistants and improve transparency and explainability by attributing sources and providing clear explanations.

Now, let us move on to a generative AI project using RAG. Imagine you're working with a massive library of documents. You need a way to quickly search and answer questions based on the content. So, manually flipping through pages takes time and effort. Wouldn't it be great to have a system that retrieves relevant information and answers your questions directly within those documents? So, that's where our Streamlit app comes in. This app utilizes the power of natural language processing and advanced retrieval techniques to turn your complex document collections into a powerful question-and-answer system.

So, let's take a look at the code behind this app. This app will allow users to ask questions about a collection of PDFs and get answers directly from the documents using the power of natural language processing.

Now, first, let's create a virtual environment. Now, in the terminal, let's type the command for setting up the environment in your editor. For that, let's type `conda create -p venv python=3.10 -y`. Okay, let's enter. In this command, the `-p venv` specifies the path and the environment name, while `-y` skips the prompts for a smoother install. Now, while that's setting up, let's create a few essential files. So, let's activate your new environment with the command `conda activate venv`. So, as you can see, our environment is ready.

Now, let's import libraries. So, let's start by importing the libraries we will need. In the first line, we will import `streamlit as st`. This gives us access to all the functionalities of Streamlit for building our web app interface. So, next, we will import `os` for various operating system functionalities. After that, now we will import libraries from `langchain`, which is a framework for building NLP pipelines. So, we will use these for tasks like text splitting, for document chain creation, prompting, retrieval, and more. So, we will explain each library in detail as we use them.

So, let's type `from langchain_groq import ChatGroq as gr` and next, we will type `from langchain.text_splitter import RecursiveCharacterTextSplitter`. So, let's type `RecursiveCharacterTextSplitter`. Again, let us type `from langchain.chains.combine_documents import create_stuff_documents_chain`. Again, `from langchain_core.prompts import ChatPromptTemplate`. Let's import `create_retrieval_chain`. Next, `import files`. `import faiss_files` from `langchain_community.vectorstores`. This will help us to create a vector index for efficient document retrieval. So, let us type `from langchain_community.vectorstores import FAISS`. Similar imports will follow for other functionalities like document loading and generating embeddings, but we will introduce them as they appear in the code.

But before this, go to the Groq Cloud website and on your left, you have the API key option. So, select and create your API key and copy this. And if you want to check your model, then go to the playground and at the top right corner, click on the Llama model and check. There are so many of them, latest also. So, choose your model and generate your free API.

Now, go to the terminal and paste it in a `.env` file using the variable `GROQ_API_KEY`. Now, again, go to the Gemini AI Studio. On your right, you have the "Create API" option. So, select your model and create your API key. Now, now copy the key and paste it into your environment variable, that is the `.env` file, using a variable `GOOGLE_API_KEY` and paste it here.

Now, we will load environment variables from a `.env` file that will securely store our API keys. So, for that, use `dotenv` to achieve this. Let's type `from dotenv import load_dotenv`. Also, let's type `load_dotenv()`. Next, we use `os` to retrieve the `GROQ_API_KEY` and `GOOGLE_API_KEY` from the environment variables using `os.getenv`. Let us type `groq_api_key = os.getenv("GROQ_API_KEY")`. And inside the function, let's type it as `google_api_key = os.getenv("GOOGLE_API_KEY")`.

So, here, these keys are required to use specific NLP services.

Now, let us write code for displaying the app title and images. So, for that, load your image. Since I'm using `edureka.png`, along with the app title "Edureka Document Question and Answer", we will use `st.image` and `st.title` for this purpose. So, for that, let us type `st.image("edureka.png", width=200)` and let us also keep the title. So, for that, `st.title("Edureka Document Question and Answers")`.

Now, the next step is to initialize ChatGroq and prompt template. Now, it's time to interact with the Langchain Groq API. So, initialize the `ChatGroq` object using `groq_api_key` and specify the `llama3-8b-8192`, which is the language model we will be using for our NLP task. So, for that, let us type `llm = ChatGroq(groq_api_key=groq_api_key, model="llama3-8b-8192")`. All right.

Now, let us define a prompt template using `ChatPromptTemplate`. So, this template ensures that AI responses are based on the context provided and user questions, keeping answers accurate and concise. So, for that, we will type `prompt = ChatPromptTemplate.from_template("""Please answer the question strictly based on the provided context. Also ensure the response is accurate, concise, and directly addresses the question. Context: {context} Question: {question}""")`.

Now, let's create a function for embedding vectors. For that, let's define `vector_embedding_function`. So, type `def vector_embedding_function():`. Next, in the next line, give `if "embeddings" not in st.session_state:`. Then type `st.session_state.embeddings = GoogleGenerativeAIEmbeddings(model="models/embedding-001")`.

Make a folder where you will load your PDF. So, I am creating `pdf_docs` and paste your PDF here. Now, set the session state loader. That is, let us type `st.session_state.loader = PyPDFDirectoryLoader("pdf_docs")`. Here, inside the double quotes, let us paste the path of the PDF. Next is the data injection. For that, let us type `st.session_state.docs = st.session_state.loader.load()`. So, this particular line of code is for data injection, and here, this particular line is for document loading. Next, let us type `st.session_state.text_splitter = RecursiveCharacterTextSplitter(chunk_size=1000, chunk_overlap=200)`. Now, here, these are for the chunk creation. Now, let us type `st.session_state.fal_documents = st.session_state.text_splitter.split_documents(st.session_state.docs)`. Let us give `[::20]`. So, this line of code is for splitting. Now, let us type `st.session_state.vectors = FAISS.from_documents(st.session_state.fal_documents, st.session_state.embeddings)`. Okay. So, this line of code is for vector OpenAI embeddings.

Now, input field for question. Let us type `prompt1`. So, give `prompt1 = st.text_input("Enter your questions from any document")`. Now, to create a button to load embeddings, let us type `if st.button("Load Edure DB"):`. So, give a colon and in the next line, let us type `vector_embedding_function()`. Next, type `st.success("Edure DB is ready for queries")`.

If `prompt1`:

`document_chain = create_stuff_documents_chain(llm, prompt)`

`retriever = st.session_state.vectors.as_retriever()`

`retrieval_chain = create_retrieval_chain(retriever, document_chain)`

`start_time = time.process_time()`

`response = retrieval_chain.invoke({"input": prompt1})`

`response_time = time.process_time() - start_time`

`st.markdown("### AI Response:")`

`st.success(f"Answer: {response['answer']}")`

`st.write(f"Response time: {response_time:.2f} seconds")`

Now, let us write code to display similar documents in an expander. So, for that, let us type `with st.expander("Document Similarity Search Results"):`. And in the next line, let us type `st.markdown("Below are the most relevant document chunks:")`. So, type `Below are the most relevant document chunks:`. Give a colon. Close the double quotes and come to the next line. Here, let us type `for i, doc in enumerate(response["context"]):`. In the next line, let us type `st.markdown(f"""<div class="card"> <p>{doc.page_content}</p></div>""", unsafe_allow_html=True)`.

So, you can also add inline styles and HTML tags, and also icons and emojis to make your application fabulous for the user.

Now, it's time for testing. For that, open your terminal and write `streamlit run your_file_name.py`. So, once you enter, and there we go. Here's our document question and answer loader. Now, select the question from the PDF you have loaded in the file and ask your loader. So, as you can see, this is my PDF. So, I'm going to copy some question from here. So, let me just copy this. Okay. Once copied, so I'm going to paste it here. So, I'm going to click on the "Load Edure DB". So, guys, as you can see, it provides an answer in context given in the PDF. So, this is our answer that it has generated. So, that's all. We have used simple Python code and Langchain techniques of RAG and some inline HTML and styles.

Have you ever wondered how massive AI models like ChatGPT are managed and optimized? That's where LLMOps, which stands for Large Language Model Operations, comes in. LLMOps is key to training, deploying, and scaling large AI models efficiently while keeping costs low and performance high. It ensures faster responses, ethical AI, and seamless integration into real-world applications.

Large Language Model Operations is a set of practices, tools, and frameworks designed to efficiently manage, deploy, and maintain large language models like ChatGPT, Claude, and Gemini in real-world applications. Just like MLOps streamlines the development of machine learning models, LLMOps optimizes the lifecycle of LLMs from data processing and training to deployment and monitoring.

Now that you know what LLMOps is, so let's explore why it's important. As LLMs become widely integrated into business applications, customer support chatbots, content generation tools, and automation systems, they need to be continuously monitored and optimized. Without proper LLMOps practices, large language models can become inefficient, leading to slower response times and increased computational costs. So, they may also become unreliable, generating outdated or biased outputs that impact user trust and decision-making. Additionally, these models can be difficult to scale and struggle to handle increasing user demand, which can result in performance bottlenecks and degraded user experience. For example, imagine running a GPT-like AI on a customer support chatbot. Without LLMOps, responses would be slow, repetitive, and expensive. LLMOps optimizes the entire workflow.

So, now that we understand why LLMOps is important, so let's take a look at how it differs from MLOps and what makes it unique. All right. So, LLMOps is a specialized branch of MLOps, but it is tailored for large-scale language models rather than traditional machine learning models. So, here are the key differences between LLMOps and MLOps. LLMOps differs from MLOps in several key aspects. So, in terms of data complexity, LLMOps require vast amounts of diverse text data, whereas MLOps typically works with structured or tabular data. Next, compute power is another major difference, as training LLMs depends on high-performance GPUs and massive cloud resources, while traditional ML models generally require lower compute power. And when it comes to real-time processing, LLMOps necessitates scalable deployment to handle continuous inference efficiently, whereas MLOps often relies on batch processing or periodic inferences. Lastly, ethical and bias considerations are more prominent in LLMOps, requiring constant monitoring to detect and mitigate biases and misleading outputs, whereas bias monitoring in MLOps is important but generally less complex compared to LLMs.

Next, let us see how LLMOps works. So, LLMOps follows a structured workflow to ensure the efficient management of large language models. So, it begins with data collection and pre-processing, where large text datasets are cleaned and structured for training. Next, model training and fine-tuning help the AI learn to understand and generate text effectively. Once trained, the model moves to deployment, where it is run on cloud servers, edge devices, or APIs for real-world applications. And during inference and optimization, the model's response speed is improved while minimizing computational cost. Next, monitoring and feedback loops play a crucial role in tracking performance and making adjustments based on real-world usage. Finally, continuous improvement ensures the model remains relevant by updating it periodically with fresh data.

And here are some real-world examples. Companies like OpenAI, Google, and Meta use LLMOps to maintain their AI products without frequent manual retraining.

So, now that we understand how LLMOps works, so let's explore some of the popular tools and frameworks that make it possible to manage and optimize large language models efficiently. LLMOps professionals rely on specialized tools to manage the model lifecycle efficiently. So, one of the top three most popular platforms is Hugging Face, an open-source tool for NLP and transformer models. MLflow is widely used for tracking experiments, model versions, and training metrics, while Kubeflow provides a scalable MLOps framework for deploying AI in Kubernetes. And companies use a combination of these tools to streamline their LLMOps pipelines and ensure smooth deployment.

Now that we have covered the tools and frameworks used in LLMOps, next, let's explore the career opportunities and the future prospects in this rapidly growing field. LLMOps is a rapidly growing field with a high demand for skilled professionals. A Machine Learning Engineer focuses on designing and optimizing LLM models, ensuring their efficiency and effectiveness. An AI Product Manager oversees AI model deployment for businesses, ensuring smooth integration into real-world applications. The role of an LLMOps Engineer involves managing AI infrastructure and scaling models for optimal performance. And if you have a background in machine learning, cloud computing, or DevOps, transitioning into LLMOps is a great move.

So, LLMOps plays a crucial role in managing large AI models efficiently, ensuring optimal performance, reduced costs, and ethical AI development. By leveraging top tools like Hugging Face, MLflow, and Kubeflow, professionals can streamline model training, deployment, and monitoring. And with the increasing adoption of AI across industries, career opportunities in LLMOps are booming, making it an exciting and rewarding field for AI enthusiasts looking to build a future in artificial intelligence. And what do you think about LLMOps? Drop your answers in the comments.

Prompt engineering is an interesting field that combines artificial intelligence and human language understanding. In this field, professionals and researchers work to create prompts or instructions that effectively guide AI systems to produce the expected outcome. Whether it's fine-tuning language models, designing prompts for specific tasks, or optimizing human-machine communication, prompt engineering is crucial for leveraging the power of AI for a variety of applications.

Imagine you're developing a virtual assistant application using a large language model such as GPT-3. The goal is to provide users with an engaging and helpful experience by designing effective prompts that generate informative and relevant responses from the model. So, let's consider a scenario in which the virtual assistant assists users with travel planning. So, here's how prompt engineering plays a major part. So, the scenario is: You're planning a trip to Paris and want the virtual assistant to provide recommendations for activities, restaurants, and landmarks to visit during your stay. So, let's say you're looking for help with a traditional prompt, and you ask: "What should I do in Paris?" And the virtual assistant will assist you like: "Here are some recommendations for activities in Paris." And here's how the enhanced prompt through prompt engineering responds to your queries. So, if you input a query that goes like: "Hey there, I'm super excited about my upcoming trip to Paris. Could you please recommend some must-visit places and activities for me?" Then the virtual assistant will generate the response as something like this: "Of course, Paris is an amazing city with so much to offer. So, here are some must-visit places and activities..." and it continues with the explanation about each place. I hope you got the idea of how an enhanced prompt provides users with an engaging and helpful experience by designing effective prompts that generate informative and relevant responses from the model.

So, now let us understand what exactly is prompt engineering. Prompt engineering is a method used in natural language processing (NLP) and machine learning. It's all about crafting clear and precise instructions to interact with large language models like GPT-3 or BERT. So, these models can generate human-like responses based on the prompts they receive. Think of prompt engineering as giving directions to these models. By crafting specific and concise prompts, we guide them to produce the response we want. So, to do this effectively, we need to understand the capabilities of the model and the problem we are trying to solve. Fine-tuning prompts allows researchers and developers to improve the performance and usability of LLMs for a variety of applications, including text generation, question answering, language translations, and others. Effective prompt engineering necessitates a thorough understanding of the underlying model's capabilities, as well as the problem domain and desired result.

Now, let's find out why prompt engineering matters for AI. So, prompt engineering is important in AI because it improves model performance, customization, and reliability. By creating clear and tailored prompts, developers can help AI models produce more accurate and relevant results, reduce biases, improve user experience, and address ethical concerns. In simple terms, prompt engineering ensures that AI systems produce useful and reliable results that meet the needs of users while adhering to ethical principles.

So, now let's consider an example in the context of text generation for generating product descriptions. Assume you're using an AI model to create product descriptions for an online store. So, without prompt engineering, you may issue a generic prompt such as: "Generate a product description for a smartphone." So, without prompt engineering, you would get something like this: "This smartphone has a high-resolution display, powerful processor, and a long-lasting battery life." The given prompt is less effective because it lacks specificity. It simply says, "Generate a product description for a smartphone." This may make it difficult to come up with an idea and write something engaging and informative. So, having a good prompt can make a significant difference in your writing. They give you a clear idea of what you need to write about and keep you focused and organized, making it easier to generate ideas and express yourself.

On the other hand, by using prompt engineering techniques, you can provide more specific instructions or constraints that will tailor the generated descriptions to the target audience or brand style. So, with prompt engineering, if you input a query such as: "Create a product description for a budget-friendly smartphone perfect for young professionals. Highlight its affordable, sleek, and packed with a top-notch camera features." And the generated response would be something like this: "Introducing our sleek and affordable smartphone designed for young professionals. With its stylish design and advanced camera features, capturing life's moments has never been easier..." and it goes on giving its key features along with it. So, through this example, we understood that prompt engineering enables the creation of a product description that is useful to the target audience and highlights specific features based on the instructions provided. So, this shows how prompt engineering can improve the importance and effectiveness of AI-generated content for specific applications.

To help AI models give accurate answers, it's important to create clear prompts. So, here are some simple rules for generating effective prompts:

First, make it clear. So, clearly explain what you want the AI to do. Unclear prompts might confuse the AI and lead to wrong answers. So, make sure that the prompt is clear. For example, the unclear prompt is something like: "Write about cars." So, where we haven't mentioned which type of car or anything much in detail. Whereas the clear prompt is: "Write a description of a red convertible sports car."

Next, give context. So, provide enough information so that the AI understands the task. So, this helps it give accurate responses that make sense in the given situation. So, for example, prompt without context is: "Write a story." Prompt with context is: "Write a story about a girl who discovers a magic book in her attic."

Next, show examples. Use examples to show the AI what you are looking for. So, this helps it understand the type of response you want. So, for example, the prompt without example is: "Describe a beach scene." So, prompts with examples is: "Describe a beach scene with palm trees, crashing waves, and people playing volleyball."

Next, keep it short. So, don't overload the AI with too much information. Short prompts help the AI focus and give quicker, more accurate responses. For example, long prompts are like this: "Write a detailed essay discussing the impact of climate change on biodiversity and ecosystems in tropical rainforests." And short prompts look something like this: "Write about climate change effects on rainforests."

Next, avoid biases. So, make sure your prompts are fair and don't include any unfair assumptions. So, biased prompts can lead to biased answers, which isn't helpful. So, for example, "Write about a woman who struggles with her weight." So, unbiased prompts are: "Write about a person overcoming challenges."

Next, set limits. So, tell the AI any rules or restrictions it needs to follow. This helps guide its response and ensures they meet your specific needs. For example, prompt without limits are: "Write a story." And prompt with limits are: "Write a story set in a haunted house with a maximum word count of 500 words." And I hope it's very clear.

Next, moving on to some examples of prompts for generating text using ChatGPT for text generation tasks. Prompts usually consist of a textual instruction or starting point that directs the model to produce coherent and relevant text. Prompts can be story prompts, questions, or incomplete sentences. Text generation prompts provide context and directions to the model, allowing it to generate human-like text responses. They influence the generated text's tone, style, and context.

So, let's say the prompt is: "Write a short story about a character who discovers a hidden treasure." So, by providing a specific storyline and theme in the prompt, the model is guided to generate a coherent and engaging narrative centered around the discovery of a hidden treasure. So, the picture illustrates how ChatGPT crafts stories with an engaging touch, making them more captivating and interesting for readers.

Next, question answering. So, prompt is: "Can you describe the common signs and symptoms of COVID-19 along with any precautions that can be taken to stay safe?" And just like that, it can generate answers to all your questions in mere seconds. So, by framing the prompt as a question, the model is directed to provide a concise answer regarding the symptoms of COVID-19, ensuring relevant and informative responses.

Next, language translation. "Translate the given English sentence 'The quick brown fox jumps over the lazy dog' into Spanish while maintaining its original meaning." So, by specifying the source and target language in the prompt, along with the input sentence, the model is instructed to perform a precise translation task, ensuring accurate language conversion.

Next, code autocompletion. Using OpenAI Codex or ChatGPT, you can perform code autocompletion tasks. So, here we go with ChatGPT. Code generation prompts are usually partial code snippets or descriptions of programming tasks. They specify the desired functionality or behavior that the model should exhibit. Code generation prompts allow the model to generate code that satisfies specific programming requirements, such as implementing algorithms, defining functions, or solving coding problems. So, the prompt is: "Complete the following Python function to calculate the factorial of a number." And here you have also added the function. So, by presenting an incomplete code snippet along with clear instructions, the model is directed to suggest appropriate code completion, helping developers write code more efficiently.

Now, moving on to text-to-image generation. Image generation prompts specify the visual scene, objects, or concepts that the model should generate. They may include textual descriptions, keywords, or images. So, image generation prompts allow the model to know what visual content to generate. They influence the generated image's composition, style, and detail. For example, the prompt is: "Imagine a tree where the branches are made of stacks of books. Can you paint me a picture of that?" And for the given prompt, we got the image generated as something like this: "An imaginative portrayal of a tree with branches composed of stacked books. Eight books representing leaves and covers visible." And the next prompt is: "Picture a cloud in the sky that looks like a huge heart. Can..."

You draw that for me? And here we go. These AI tools leverage prompt engineering techniques to generate text, perform language translation, code auto-completion, and text-to-image generation, demonstrating the versatility and power of prompt-based interactions with AI models.

Next, why is machine learning useful in prompt engineering? Machine learning is very helpful in prompt engineering, especially in linguistic and language models, because it helps create better prompts and interactions by analyzing lots of data and finding patterns.

So, first, understanding language patterns. Machine learning algorithms can analyze large amounts of text to understand linguistic patterns like grammar, syntax, semantics, and context. So, this understanding is critical for developing effective prompts that generate desired responses from language models.

Next, generating relevant prompts. Machine learning models can suggest or generate prompts based on input data and user preferences. These prompts can be tailored to specific tasks, domains, or user requirements, making them more useful and efficient for guiding language models.

Next, optimizing prompt design. Machine learning techniques can be used to optimize prompt design by comparing the performance of various prompts and selecting the one that produces the best result. This iterative process improves prompt engineering practices and the overall performance of language models.

And the next is personalizing interactions. Machine learning enables personalized interactions by creating prompts tailored to individual users' preferences, history, and context. This personalization increases user engagement and satisfaction with the language model interaction.

Next, improving model performance. Machine learning algorithms can be used to fine-tune language models based on prompt-response pairs, increasing their performance and accuracy over time. Language models can be trained on a variety of datasets and prompts to produce more relevant and contextually appropriate responses.

And next, mitigating bias and misinformation. Machine learning techniques can help identify and mitigate biases in prompt engineering by examining prompt-response pairs for potential biases or inaccuracies. Language models can produce more fair, inclusive, and reliable results by detecting and correcting for biases.

And I hope it is clear why machine learning is useful in prompt engineering. Now, the success of the human race is because of the ability to communicate and share information. Now, that is where the concept of language comes in. However, many such standards came up, resulting in many such languages, with each language having its own set of basic shapes called alphabets, and the combination of alphabets resulted in words, and the combination of these words, arranged meaningfully, resulted in the formation of a sentence. Now, each language has a set of rules that is used while developing these sentences, and these sets of rules are also known as grammar.

Now, coming to today's world, that is the 21st century. According to industry estimates, only 21% of the available data is present in a structured format. Data is being generated as we speak, as we tweet, as we send messages on WhatsApp, Facebook, Instagram, or through text messages, and the majority of this data exists in the textual form, which is highly unstructured in nature.

Now, in order to produce significant and actionable insights from text data, it is important to get acquainted with the techniques of text analysis. So, let's understand what is text analysis or text mining. Now, it is the process of deriving meaningful information from natural language text, and text mining usually involves the process of structuring the input text, deriving patterns within the structured data, and finally evaluating the interpreted output.

Compared with the kind of data stored in databases, text is unstructured, amorphous, and difficult to deal with algorithmically. Nevertheless, in modern culture, text is the most common vehicle for the exchange of information.

Now, as text mining refers to the process of deriving high-quality information from text, the overall goal here is to turn the text into data for analysis, and this is done by the application of NLP, or natural language processing.

So, let's understand what is natural language processing. So, NLP refers to the artificial intelligence method of communicating with an intelligence system using natural language. By utilizing NLP and its components, one can organize massive chunks of textual data, perform numerous automated tasks, and solve a wide range of problems such as automatic summarization, machine translation, named entity recognition, speech recognition, and topic segmentation.

So, let's understand the basic structure of an NLP application. Considering the chatbot here as an example, we can see first we have the NLP layer, which is connected to the knowledge base and the data storage. Now, the knowledge base is where we have the source content, that is, we have all the chat logs which contain a large history of all the chats which are used to train the particular algorithm. And again, we have the data storage where we have the interaction history and the analytics of that interaction, which in turn helps the NLP layer to generate the meaningful output.

So, now, if we have a look at the various applications of NLP. First of all, we have sentiment analysis. Now, this is a field where NLP is used heavily. We have speech recognition. Now, here we are also talking about the voice assistants like Google Assistant, Cortana, and Siri. Now, next, we have the implementation of chatbots, as I discussed earlier, just now. Now, you might have used the customer care chat services of any app. It also uses NLP to process the data entered and provide the response based on the input.

Now, machine translation is also another use case of natural language processing. Now, considering the most common example here would be Google Translate. It uses NLP and translates the data from one language to another, and that too in real-time.

Now, other applications of NLP include spell-checking. Then we have keyword search, which is also a big field where NLP is used. Extracting information from any particular website or any particular document is also a use case of NLP. And one of the coolest applications of NLP is advertisement matching. Now, here what we mean is basically recommendation of ads based on your history.

Now, NLP is divided into two major components: that is, the natural language understanding, which is also known as NLU, and we have the natural language generation, which is also known as NLG. The understanding involves tasks like mapping the given input into natural language into useful representations, analyzing different aspects of the language. Whereas natural language generation, it is the process of producing meaningful phrases and sentences in the form of natural language. It involves text planning, sentence planning, and text realization.

Now, NLU is usually considered harder than NLG. Now, you might be thinking that even a small child can understand a language. So, let's see what are the difficulties a machine faces while understanding any particular language.

Now, understanding a new language is very hard. Taking our English into consideration, there are a lot of ambiguities, and that too at different levels. We have lexical ambiguity, syntactical ambiguity, and referential ambiguity.

So, lexical ambiguity is the presence of two or more possible meanings within a single word. It is also sometimes referred to as semantic ambiguity. For example, let's consider these sentences and let's focus on the italicized words. "She is looking for a match." So, what do you infer by the word "match"? Is it that she is looking for a partner, or is it that she's looking for a match, be it a cricket match or a rugby match?

Now, the second sentence here: "The fisherman went to the bank." Is it the bank where we go to collect our checks and money, or is it the river bank we are talking about here? Sometimes it is obvious that we are talking about the river bank, but it might be true that he's actually going to a bank to withdraw some money. You never know.

Now, coming to the second type of ambiguity, which is syntactical ambiguity. In English grammar, this syntactical ambiguity is the presence of two or more possible meanings within a single sentence or a sequence of words. It is also called as structural ambiguity or grammatical ambiguity. Taking these sentences into consideration, we can clearly see what are the ambiguities faced. "The chicken is ready to eat." So, here what do you infer? Is the chicken ready to eat its food, or is the chicken ready for us to eat? Similarly, we have the sentence like "Visiting relatives can be boring." Are the relatives boring, or when we are visiting the relatives, it is very boring? You never know.

Coming to the final ambiguity, which is referential ambiguity. Now, this ambiguity arises when we are referring to something using pronouns. "The boy told his father the theft. He was very upset." Now, I'm leaving this up to you. You tell me what does "he" stand for here? Who is he? Is it the boy? Is it the father, or is it the thief?

So, coming back to NLP. Firstly, we need to install the NLTK library, that is, the Natural Language Toolkit. It is the leading platform for building Python programs to work with human language data, and it also provides easy-to-use interfaces to work with 15 corpora and lexical resources. We can use it to perform functions like classification, tokenization, stemming, tagging, and much more.

Now, once you install the NLTK library, you will see an NLTK downloader. It is a pop-up window which will come up, and in that, you have to select the "all" option and press the download button. It will download all the required files, the corpora, the models, and all the different packages which are available in the NLTK.

Now, when we process text, there are a few terminologies that we need to understand. Now, the first one is tokenization. So, tokenization is a process of breaking strings into tokens, which in turn are small structures or units that can be used for tokenization. Now, tokenization involves three steps: which is the breaking a complex sentence into words, understanding the importance of each word with respect to the sentence, and finally, producing a structural description of an input sentence.

So, if we have a look at the example here, considering this sentence: "Tokenization is the first step in NLP." Now, when we divide it into tokens, as you can see here, we have 1, 2, 3, 4, 5, 6, and 7 tokens here.

Now, NLTK also allows you to tokenize phrases containing more than one word. So, let's go ahead and see how we can implement tokenization using NLTK. So, here I'm using Jupyter Notebook to execute all my practicals and demos. Now, you are free to use any sort of IDE which is supported by Python. It's your choice. So, let me create a new notebook here. Let me rename it as "Text Mining and NLP."

So, first of all, let us import all the necessary libraries. Here we are importing the OS, NLTK, and the NLTK corpus. So, as you can see here, we have various files which represent different types of words, different types of functions. We have samples of Twitter, we have different sentimental word nets, we have product reviews, we have movie reviews, we have non-breaking prefixes, and many more files here.

Now, let's have a look at the Gutenberg file here and see what are all the fields which are present in the Gutenberg file. So, as you can see here, inside this, we have all the different types of text files. We have Austin's Emma, we have Shakespeare, we have Hamlet, we have Moby Dick, we have Carroll's Alice, and many more. Now, this is just one file we are talking about, and NLTK provides a lot of files.

So, let's consider a document of type string and understand the significance of its tokens. So, if you have a look at the elements of Hamlet, you can see it starts from "The Tragedy of Hamlet by William Shakespeare, 1599, Actor's Premise." We can use a lot of these files for analysis and text for understanding and analysis purposes, and this is where NLTK comes into the picture, and it helps a lot of programmers to learn about the different features and the different applications of language processing.

So, here I have created a paragraph on artificial intelligence. So, let me just execute it. Now, this AI is of the string type, so it will be easier for us to tokenize it. Nonetheless, any of the files can be used to tokenize. For simplicity, here I'm taking a string file.

The next what we are going to do is import the word_tokenize under the NLTK.tokenize library. Now, this will help us to tokenize all the words. Now, we will run the word_tokenize function over the paragraph and assign it a name. So, here I'm considering AI_tokens and I'm using the word_tokenize function on it. Let's see what's the output of this AI_tokens. So, as you can see here, it has divided all the input which was provided here into the tokens.

Now, let's have a look at the number of tokens here we have here. So, in total, we have 273 tokens. Now, these tokens are a list of words and the special characters which are separated items of the list.

Now, in order to find the frequency of the distinct elements here in the given AI paragraph, we are going to import the FreqDist function which falls under nltk.probability. So, let's create a FreqDist in which we have the function here, FreqDist, and basically what we are doing here is finding the word count of all the words in the paragraph. So, as you can see here, we have comma 30 times, we have full stop nine times, and we have accomplished one, one, and so on. We have computer five times.

Now, here we are also converting the tokens into lowercase so as to avoid the possibility of considering a word with uppercase and lowercase as different. Now, suppose we were to select the top 10 tokens with the highest frequency. So, here you can see that we have comma 30 times, the 13 times, of 12 times, and 12 times, whereas the meaningful words which are intelligence, which is six times, and intelligence six times.

Now, there is another type of tokenizer, which is the blank tokenizer. Now, let's use the blank tokenizer over the same string to tokenize the paragraph with respect to the blank string. Now, the output here is nine. Now, this nine indicates how many paragraphs we have and what all paragraphs are separated by a new line. Although it might seem like one paragraph, it is not. The original structure of the data remains intact.

Now, another important key term in tokenization are bigrams, trigrams, and n-grams. Now, what does this mean? Now, bigrams refers to tokens of two consecutive words, known as a bigram. Similarly, tokens of three consecutive written words are known as trigrams. And similarly, we have n-grams for the n consecutive written words.

So, let's go ahead and execute some demos based on bigrams, trigrams, and n-grams. So, first of all, what we need to do is import bigrams, trigrams, and n-grams from nltk.util. Now, let's take a string here on which we'll use these functions. So, taking this string into consideration: "The best and the most beautiful thing in the world cannot be seen or even touched. They must be felt with the heart."

So, first, what we are going to do is split the above sentence or the string into tokens. So, for that, we are going to use the word_tokenize. So, as you can see here, we have the tokens. Now, let us now create the bigrams of the list containing tokens. So, for that, we are going to use the nltk.bigrams and pass all the tokens, and since it is a list, we are going to use the list function. So, as you can see under output, we have "the best," "best and," "and the," "the most," "most beautiful," "beautiful thing," "thing in," "in the," "the world." So, as you can see, the tokens are in the form of two words; it's in a pair form.

Similarly, if we want to do the trigrams and find out the trigrams, what we need to do is just remove the bigrams and use the trigrams. So, as you can see, we have tokens in the form of three words. And if you want to use the n-grams, let me show you how it's done. So, for n-grams, what we need to do is define a particular number here. So, instead of n, I'm going to use, let's say, four. So, as you can see, we have the output in the form of four tokens.

Now, once we have the tokens, we need to make some changes to the tokens. So, for that, we have stemming. Now, stemming usually refers to normalizing words into its base form or the root form. So, if we have a look at the words here: "affectation, effects, affections, affected, affection, and affecting." So, as you might have guessed, the root word here is "affect." So, one thing to keep in mind here is that the result may not be the root word always. The stemming algorithm works by cutting off the end or the beginning of the word, taking into account a list of common prefixes and suffixes that can be found in an inflected word. Now, this indiscriminate cutting can be successful in some occasions, but not always, and this is why we affirm that this approach presents some limitations.

So, let's go ahead and see how we can perform stemming on a particular given dataset. Now, there are quite a few types of stemmers. So, starting with the Porter stemmer, we need to import it from nltk.stem. Let's get the output of the word "having" and see what is the stemming of this word. So, as you can see, we have "hav" as the output. Now, here we have defined words to stem, which are "give," "giving," "given," and "gave." So, let's use the Porter stemmer and see what is the output of this particular stemming. So, as you can see, it has given "give," "given," "give," and "gave." Now, we can see that the stemmer removed only the "ing" and replaced it with an "e."

Now, let's try to do it the same with another stemmer called the Lancaster stemmer. You can see the stemmer stemmed all the words. As a result of it, you can conclude that the Lancaster stemmer is more aggressive than the Porter stemmer. Now, the use of each of these stemmers depends on the type of task that you want to perform. For example, if you want to check how many times the words "GIV" is used above, you can use the Lancaster stemmer. And for other purposes, you have the Porter stemmer as well. Now, there are a lot of stemmers. There is one Snowball stemmer also present, where you need to specify the language which you are using and then use the Snowball stemmer.

Now, as we discussed that the stemming algorithm works by cutting off the end or the beginning of the word. On the other hand, lemmatization takes into consideration the morphological analysis of the word. Now, in order to do so, it is necessary to have a detailed dictionary which the algorithm can look into to link the form back to its lemma. Now, lemmatization, what it does is groups together different inflected forms of a word, which are called lemmas. It is somehow similar to stemming, as it maps several words into a common root.

Now, one of the most important things here to consider is that the output of lemmatization is a proper word, unlike stemming, in that case, where we got the output as "GIV." Now, "GIV" is not any word; it's just a stem. Now, for example, if a lemmatizer should work on "go," "going," and "went," it all stems into "go" because that is the root of all the three words here.

So, let's go ahead and see how lemmatization works on the given input data. Now, for that, we are going to import the lemmatizer from NLTK. Now, we are also importing the WordNet here. As I mentioned earlier, that lemmatization requires a detailed dictionary because the output of it is a root word, which is a particular given word. It's not just any random word; it is a proper word. So, to find that proper word, it needs a dictionary. So, here we are providing the WordNet dictionary and we are using the WordNet lemmatizer. So, passing the word "corpora" into the WordNet lemmatizer. So, can you guys tell me what is the output of this one? I'll leave this up to you guys. I won't execute this sentence. Let me remove this sentence here. You guys tell me in the comments below what will be the output of the lemmatization of the word "corpora." And what will be the output of the stemming? You guys execute that and let me know in the comment section below.

Now, let's take these words into consideration: "give," "giving," "given," and "gave," and see what is the output of the lemmatization. So, as you can see here, the lemmatizer has kept the words as it is, and this is because we haven't assigned any POS tags here, and hence it has assumed all the words as nouns. Now, you might be wondering what are POS tags? Well, I'll tell you what are POS tags later in this video. So, for just now, let's keep it as simple as that is that POS tags usually tell us what exactly the given word is. Is it a noun? Is it a verb, or is it a different part of speech? Basically, POS stands for Parts of Speech.

Now, do you know that there are several words in the English language such as "I," "ate," "for," "above," "below," which are very useful in the formation of sentences, and without them, the sentence wouldn't make any sense. But these words do not provide any help in natural language processing, and this list of words are also known as stop words. NLTK has its own list of stop words, and you can use the same by importing it from the NLTK.corpus.

So, the question arises, are they helpful or not? Yes, they are helpful in the creation of sentences, but they are not helpful in the processing of the language. So, let's check the list of stop words in the NLTK. So, from NLTK.corpus, we are importing the stop words, and if we specify what all stop words are there in the English language, let's see. So, as you can see here, we have the list of all the stop words which are defined in the English language, and we have 179 total number of stop words.

Now, as you can see here, we have these words which are "few," "more," "most," "other," "some." Now, these words are very necessary in the formation of sentences. You cannot ignore these words, but for processing, these are not important at all. So, if you remember, we had the top 10 tokens from that particular word, that is, the AI paragraph I mentioned earlier, which was given as FreqDist top 10. Let's take that into consideration and see what we can see here is that except "intelligent" and "intelligence," most of the words are either punctuation or stop words, and hence can be removed.

Now, we'll use the `compile` from the `re` module to create a string that matches any digit or special character, and then we'll see how we can remove the stop words. So, if you have a look at the output of the post-punctuation, you can see there are no stop words here in the particular given output. And if you have a look at the output of the length of the post-punctuation, it's 233 compared to the 273, the length of the AI tokens. Now, this is very necessary in language processing as it removes all the unnecessary words which do not hold any much more meaning.

Now, coming to another important topic of natural language processing and text mining or text analysis is the parts of speech. Now, generally speaking, the grammatical type of the word, which is the verb, noun, adjective, adverb, article, indicates how a word functions in the meaning as well as the grammatical structure within the sentence. Now, a word can have more than one part of speech based on the context in which it is used. For example, if we take the sentence into consideration: "Google something on the internet." Now, here, "Google" acts as a verb, although it is a proper noun.

So, as you can see here, we have so many types of POS tags, and we have the descriptions of those various tags. So, we have the coordinating conjunction (CC), cardinal number (CD), we have JJ as adjective, MD as modal, we have the proper noun singular (NNP), plural (NNPS), we have verbs (different types of verbs), we have interjection (UH), symbol (SYM), we have the PRP pronoun, and the RB adverb.

Now, we can use POS tags as a statistical NLP task. It distinguishes the sense of the word, which is very helpful in text realization, and it is easy to evaluate as in how many tags are correct, and you can also infer semantic information from the given text.

So, let's have a look at some of the examples of POS. So, take the sentence: "The dog killed the bat." So, here, "the" is a determiner, "dog" is a noun, "killed" is a verb, and again, "the" and "bat" are determiner and noun respectively. Now, let's consider another sentence: "The way to clear the plates from the table." So, as you can see here, all the tokens here correspond to a particular type of tag, which is the part of speech tag. It is very helpful in text realization.

Now, let's consider a string and check how NLTK performs POS tagging on it. So, let's take the sentence: "Timothy is a natural when it comes to drawing." First, we are going to tokenize it. And under NLTK only, we have the POS tag option, and we'll pass all the tokens here. So, as you can see, we have "Timothy" as noun (NN), "is" as verb (VBZ), "a" as determiner (DT), "natural" as an adjective (JJ), "when" as a conjunction (WRB), "it" as a pronoun (PRP), "comes" as a verb (VBZ), "to" as a preposition (TO), and "drawing" as a verb again (VBG). So, this is how you define the POS tags. The POS tag function does all the work here.

Now, let's take another example here: "John is eating a delicious cake." And let's see what's the output of this one. Now, here you can see that the tagger has tagged both the word "is" and "eating" as a verb because it has considered "is eating" as a single term. This is one of the few shortcomings of the POS taggers. One thing important to keep in mind.

Now, after POS tagging, there is another important topic, which is the named entity recognition. So, what does it mean? Now, the process of detecting the named entities such as the person name, the location name, the company name, the organization, the quantities, and the monetary value is called the named entity recognition. Now, under named entity recognition, we have three types of identification. Here we have the noun phrase identification. Now, this step deals with extracting all the noun phrases from a text using dependency parsing and parts of speech tagging. Then we have the phrase classification. The step classification. This is the classification step in which all the extracted noun phrases are classified into respective categories, which are the location, names, organization, and much more. And apart from this, one can curate the lookup tables and dictionaries by combining information from different sources. And finally, we have the entity disambiguation. Now, sometimes it is possible that the entities are misclassified. Hence, creating a validation layer on top of the result is very useful, and the use of knowledge graphs can be exploited for this purpose. Now, the popular knowledge graphs are Google Knowledge Graph, the IBM Watson, and Wikipedia.

So, let's take a sentence into consideration: "The Google CEO Sundar Pichai introduced the new Pixel at Minnesota Roy Center event." So, as you can see here, "Google" is an organization (ORG), "Sundar Pichai" as a person (PERSON), "Minnesota" is a location (GPE), and the "Roy Center event" is also tagged as an organization (ORG).

Now, for using in Python, we'll have to import the `ne_chunk` from the NLTK module, which is present in Python. So, let's consider a text data here and see how we can perform the NE using the NLTK library. So, first, we need to import the `ne_chunk` here. Let's consider the sentence here: "The US president stays in the White House." So, we need to do all these processes again. We need to tokenize the sentence first and then add the POS tags. And then, if we use the `ne_chunk` function and pass the list of tuples containing POS tags to it, let's see the output. So, as you can see, "The US" here is recognized as an organization (ORGANIZATION), and "White House" is clubbed together as a single entity and is recognized as a facility (FACILITY). Now, this is only possible because of the POS tagging. Without the POS tagging, it would be very hard to detect the named entities of the given tokens.

Now that we have understood what are named entity recognition, and yes, let's go ahead and understand one of the most important topics in NLP and text mining, which is syntax. So, what is syntax? So, in linguistics, syntax is the set of rules, principles, and the processes that govern the structure of a given sentence in a given language. The term syntax is also used to refer to the study of such principles and processes. So, what we have here are certain rules as to what part of the sentence should come at what position. With these rules, one can create a syntax tree whenever there is a sentence input.

Now, a syntax tree, in layman terms, is basically a tree representation of the syntactic structure of the sentence of the strings. It is a way of representing the syntax of a programming language as a hierarchical tree structure. This structure is used for generating symbol tables for compilers and later code generation. The tree represents all the constructs in the language and their subsequent rules.

So, let's consider the statement: "The cat sat on the mat." So, as you can see here, the input is a sentence or a verb phrase, and it has been classified into a noun phrase (NP). Then the prepositional phrase (PP). Again, the noun phrase (NP) is classified into article (DT) and noun (NN), and again we have the verb (VB) which is "sat." And finally, we have the preposition (IN), "on," the article (DT), "the," and the noun (NN), which are "mat."

In order to render syntax trees in our notebook, you need to install Ghostscript, which is a rendering engine. Now, this takes a lot of time, and let me show you from where you can download Ghostscript. Just type in "download Ghostscript" and select the latest version here. So, as you can see, we have two types of licenses here. We have the General Public License and the commercial license. As creating syntax and following it is a very important part, it is also available for commercial license, and it is very useful. So, I'm not going to go much deeper into what a syntax tree is and how we can do that.

So, now that we have understood what are syntax trees, let's discuss the important concept with respect to analyzing the sentence structure, which is chunking. So, chunking basically means picking up individual pieces of information and grouping them into bigger pieces. And these bigger pieces are also known as chunks. In the context of NLP and text mining, chunking means grouping of words or tokens into chunks.

So, let's have a look at the example here. So, the sentence into consideration here is: "We caught the black panther." "We" is a pronoun (PRP). "caught" is a verb (VBD). "the" is a determiner (DT). "black" is an adjective (JJ), and "panther" is a noun (NN). So, what it has done is here, as you can see, is that "black" (JJ), "panther" (NN), and "the" (DT) are chunked together in the noun phrase (NP).

So, let's go ahead and see how we can implement chunking using the NLTK. So, let's take the sentence: "The big cat ate little mouse who was after the fresh cheese." We'll use the POS tags here and also use the tokenizing function here. So, as you can see here, we have the tokens and we have the POS tags. What we'll do now is create a grammar from a noun phrase, and we'll mention the tags that we want in our chunk phrase within the curly braces. So, that will be our grammar `NP`. Now, here we have created a regular expression matching string. Now, we'll now have to pass the chunk, and hence we'll create a chunk parser and pass our noun phrase string to it. So, as you can see, we have a certain error, and let me tell you why this error occurred. So, this error occurred because we did not use the Ghostscript and we did not form the syntactical tree. But in the final output, we have a tree structure here, which is not exactly in the visualization part, but it's there. So, as you can see here, we have the NP (noun phrase) for "the little mouse." Again, we have the noun phrase (NP) for "fresh cheese" also. Although "fresh" is an adjective and "cheese" is a noun, it has considered a noun phrase of these two words. So, this is how you execute chunking in the NLTK library.

So, by now, we have learned almost all the important steps in text processing, and let's apply them all in building a machine learning classifier on the movie reviews from the NLTK corpora. For that, first, let me import all the libraries, which are the pandas, the numpy library. Now, these are the basic libraries needed in any machine learning algorithm. We are also importing the CountVectorizer. I'll tell you why it is used later. Now, let's just import it for now.

So, again, if we have a look at the different elements of the corpora, as we saw earlier in the beginning of our session, we have so many files in the given NLTK corpora. Now, let's now access the movie reviews corpora under the NLTK corpora. As you can see here, we have the movie reviews. So, for that, we are going to import the movie reviews from the NLTK corpus. So, if you have a look at the different categories of the movie reviews, we have two categories, which are the negative and the positive. So, if you have a look at the positive, we can see we have so many text files here. Similarly, if we have a look at the negative, we have a thousand negative files also here, which have the negative feedbacks.

So, let's take a particular positive one into consideration, which is the `cv0029590`. You can take any one of the files here; it doesn't matter. Now, the above tokenization, as you can see here, the file is already tokenized, but it is generally useful for us to do the tokenization, but the above tokenization has increased our work here. And in order to use the CountVectorizer and the TF-IDF, we must pass the strings instead of the tokens. Now, in order to convert the strings into tokens, we can use the `word_tokenize` within the NLTK, but that has some licensing issues as of now with the environment. So, instead of that, we can also use the `join` method to join all the tokens of the list into a single string. And that's what we are going to use here.

So, first, we are going to create an empty list and append all the tokens within it. We have the `review_list` that is an empty list. Now, what we are going to do here is remove all the extra spaces, the commas from the list while appending it to the empty list, and perform the same for the positive and the negative reviews. So, this one we are doing it for the negative reviews, and then we'll do the same for the positive reviews as well. So, if you have a look at the length of this negative review list, it's 1,000. And the moment we add the positive reviews also, I think the length should reach 2,000. So, let me just define the positive reviews. Now, execute the same for positive reviews. And then again, if you have a look at the length of the review list, it should be 2,000. That is good.

Now, let us now create the targets before creating the features for our classifiers. So, while creating the targets, we are using the negative reviews here, we are denoting it as zero, and for the positive reviews, we are converting it into one, and also we will create an empty list and we'll add 1,000 zeros followed by 1,000 ones into the empty list. Now, we'll create a pandas Series for the target list. Now, the type of `y` must result into a pandas Series. So, if we have a look at the output of the type of `y`, it is `pandas.core.series.Series`. That is good. Now, let's have a look at the first five entries of the Series. So, as you can see, it is 1,000 zeros, which will be followed by 1,000 ones. So, the first five inputs are all zeros.

Now, we can start creating features using the CountVectorizer or the bag of words. For that, we need to import the CountVectorizer. Now, once we have initialized the vectorizer, now we need to fit it onto the `rev_list`. Now, let us now have a look at the dimensions of this particular vector. So, as you can see, it's 2,000 by 16,228. Now, we are going to create a list with the names of all the features by typing the vectorizer name. So, as you can see here, we have our list. Now, what we'll do is we'll create a pandas DataFrame by passing the SciPy matrix as values and feature names as the column names. Now, let us now check the dimension of this particular pandas DataFrame. So, as you can see, it's the same dimension, 2,000 by 16,228. Now, if we have a look at the top five rows of the DataFrame. So, as you can see here, we have 16,228 columns with five rows, and all the inputs are here, up to zero.

Now, the DataFrame we are going to do is now split it into training and testing sets, and let us now examine the training and the test sets as well. So, as you can see, the size here we have defined as 0.25, that is, the test set, that is 25%; the training set will have the 75% of the particular DataFrame. So, if you have a look at the shape of `X_train`, we have 15,000. And if you have a look at the dimension of `X_test`, this is 5,000. So, now our data is split.

Now, we'll use the Naive Bayes classifier for text classification over the training and testing sets. So, now, most of you guys might already be aware of what a Naive Bayes classifier is. So, it is basically a classification technique based on the Bayes' theorem with an assumption of independence among predictors. In simple terms, a Naive Bayes classifier assumes that the presence of a particular feature in a class is unrelated to the presence of any other feature. To know more, you can watch our Naive Bayes classifier video, the link to which is given in the description box below. If you want to pause at this moment of time and check quickly what a Naive Bayes classifier does and how it works, you can check that video and come back here.

Now, to implement the Naive Bayes algorithm in Python, we'll use the following library and the functions. We are going to import the `GaussianNB` from the scikit-learn library, which is scikit-learn. We are going to instantiate the classifier now and fit the classifier with the training features and the labels. We are also going to import the `MultinomialNB` because we do not have only two features here; we have the multinomial features. So, now we have passed the training and the test dataset to this particular Multinomial Naive Bayes, and then we will use the `predict` function and pass the training features.

Now, let's have a look and check the accuracy of this particular metrics. So, as you can see here, the accuracy here is one, that is very highly unlikely, but since it has given one, that means it is overfitting and it is overly accurate, and you can also check the confusion matrix for the same. For that, what you need to do is use the confusion matrix on these variables, which is `y_test` and `y_predicted`. So, as you can see here, although it has predicted 100% accuracy, the accuracy is one. This is very highly unlikely, and you might have got a different output for this one. I've got the output here as 1.0. You might have got an output as 0.6, 0.7, or any number in between 0 and 1.

What is the KNN algorithm? Well, K-Nearest Neighbor is a simple algorithm that stores all the available cases and classifies the new data or case based on a similarity measure. It suggests that if you are similar to your neighbors, then you are one of them. Right? For example, if an apple looks more similar to a banana, orange, or melon rather than a monkey, rat, or a cat, then most likely an apple belongs to the group of fruits. All right? Well, in general, KNN is used in search applications where you're looking for similar items. That is, when your task is some form of finding items similar to this one, then you call this search as a KNN search. But what is this K in KNN? Well, the K denotes the number of nearest neighbors which are voting for the class of the new data or the testing data. For example, if K=1, then the testing data is given the same label as the closest example in the training set. Similarly, if K=3, the labels of the three closest classes are checked, and the most common label is assigned to the testing data. So, this is what a KNN algorithm means.

So, moving on ahead, let's see some of the examples of scenarios where KNN is used in the industry. So, let's see the industrial application of KNN algorithm. Starting with recommendation systems. Well, the biggest use case of KNN search is a recommendation system. This recommendation system is like an automated form of a shop counter guide. When you ask him for a product, not only does he show you the product, but also suggests you or displays you a relevant set of products which are related to the item you're already interested in buying. This KNN algorithm applies to recommending products like in Amazon, or for recommending media like in the case of Netflix, or even for recommending advertisements to display to a user. If I'm not wrong, almost all of you must have used Amazon for shopping, right? So, just to tell you, more than 35% of Amazon.com's revenue is generated by its recommendation engine. So, what's their strategy? Amazon uses recommendations as a targeted marketing tool in both the email campaigns and on most of its website pages. Amazon will recommend many products from different categories based on what you are browsing, and it will pull those products in front of you which you're likely to buy, like the "frequently bought together" option that comes at the bottom of the product page to tempt you into buying the combo. Well, this recommendation has just one main goal: that is, to increase the average order value or to upsell and cross-sell customers by providing product suggestions based on items in the shopping cart or based on the product they are currently looking at on the site.

So, next industrial application of KNN algorithm is concept search, or searching semantically similar documents and classifying documents containing similar topics. As you know, the data on the internet is increasing exponentially every single second. There are billions and billions of documents on the internet. Each document on the internet contains multiple concepts that could be a potential concept. Now, this is a situation where the main problem is to extract concepts from a set of documents, as each page could have thousands of combinations that could be potential concepts. An average document could have millions of concepts. Combine that with the vast amount of data on the web. Well, we are talking about an enormous amount of dataset and samples. So, what we need here? We need to find a concept from the enormous amount of dataset and samples, right? So, for this purpose, we'll be using the KNN algorithm. More advanced examples could include handwriting detection like an OCR, or image recognition, or even video recognition. All right?

So, now that you know various use cases of KNN algorithm, let's proceed and see how does it work. So, how does a KNN algorithm work? Let's start by plotting these blue and orange points on our graph. So, these blue points, they belong to class A, and the orange ones, they belong to class B. Now, you get a star as a new point, and your task is to predict whether this new point belongs to class A or it belongs to class B.

So, to start the prediction, the very first thing that you have to do is select the value of K. Just as I told you, K in KNN algorithm refers to the number of nearest neighbors that you want to select. For example, in this case, K=3. So, what does it mean? It means that I'm selecting three points which are the least distance to the new point, or you can say I'm selecting three different points which are closest to the star. Well, at this point of time, you can ask, how will you calculate the least distance? So, once you calculate the distance, you'll get one blue and two orange points which are closest to the star. Now, since in this case, as we have a majority of orange points, so you can say that for K=3, the star belongs to class B, or you can say that the star is more similar to the orange points.

Moving on ahead. Well, what if K=6? Well, for this case, you have to look for six different points which are closest to the star. So, in this case, after calculating the distance, we find that we have four blue points and two orange points which are closest to the star. Now, as you can see that the blue points are in majority, so you can say that for K=6, the star belongs to class A, or the star is more similar to blue points. So, by

Now, I guess you know how a KNN algorithm works and what is the significance of K in the KNN algorithm. So, how will you choose the value of K? So, keeping in mind, this K is the most important parameter in the KNN algorithm. So, let's see when you build a K-nearest neighbor classifier, how will you choose a value of K? Well, you might have a specific value of K in mind, or you could divide up your data and use something like a cross-validation technique to test several values of K in order to determine which works best for your data. For example, if N equals a thousand cases, then in that case, the optimal value of K lies somewhere in between 1 to 19. But yes, unless you try it, you cannot be sure of it.

So, you know how the algorithm is working on a higher level. Let's move on and see how things are predicted using the KNN algorithm. Remember, I told you the KNN algorithm uses the least distance measure in order to find its nearest neighbors. So, let's see how these distances are calculated. Well, there are several distance measures which can be used. So, to start with, we'll mainly focus on Euclidean distance and Manhattan distance in this session.

So, what is this Euclidean distance? Well, this Euclidean distance is defined as the square root of the sum of the difference between a new point X and an existing point Y. So, for example, here we have point P1 and P2. Point P1 is (1,1,1) and point P2 is (5,4). So, what is the Euclidean distance between both of them? So, you can say that Euclidean distance is the direct distance between two points. So, what is the distance between the point P1 and P2? So, we can calculate it as (5-1)^2 + (4-1)^2, and we can root it over, which results to 5.

So, next is the Manhattan distance. Well, this Manhattan distance is used to calculate the distance between real vectors using the sum of their absolute differences. In this case, the Manhattan distance between the point P1 and P2 is |5-1| + |4-1|, which results to 3 + 4, that is 7. So, this slide shows the difference between Euclidean and Manhattan distance from point A to point B. So, Euclidean distance is nothing but the direct or the least possible distance between A and B. Whereas the Manhattan distance is the distance between A and B measured along the axes at right angles.

Let's take an example and see how things are predicted using the KNN algorithm or how the KNN algorithm is working. Suppose we have a dataset which consists of height, weight, and t-shirt size of some customers. Now, when a new customer comes, we only have his height and weight as the information. Now, our task is to predict what is the t-shirt size of that particular customer. So, for this, we'll be using the KNN algorithm. So, the very first thing what we need to do is we need to calculate the Euclidean distance. So, now that you have new data of height 161 cm and weight as 61 kg. So, the very first thing that we'll do is we'll calculate the Euclidean distance, which is nothing but the square root of (161-158)^2 + (61-58)^2, and the square root of that is 4.24.

Let's drag and drop it. So, these are the various Euclidean distances of other points. Now, let's suppose K = 5. Then, the algorithm, what it does? It searches for the five customers closest to the new customer, that is, most similar to the new data in terms of its attributes. For K = 5, let's find the top five minimum Euclidean distances. So, these are the distances which we are going to use: 1, 2, 3, 4, and 5. So, let's rank them in order. First, this is second. This is third. Then, this one is fourth. And again, this one is five. So, this is our order. So, for K = 5, we have four t-shirts which come under size M and one t-shirt which comes under size L. So, obviously, the best guess or the best prediction for the t-shirt size of height 161 cm and weight 61 kg is M, or you can say that our new customer fits into size M.

Well, this was all about the theoretical session. But before we drill down to the coding part, let me just tell you why people call KNN as a lazy learner. Well, KNN for classification is a very simple algorithm, but that's not why they are called lazy. KNN is a lazy learner because it doesn't have a discriminative function from the training data. But what it does, it memorizes the training data. There is no learning phase of the model, and all of the work happens at the time a prediction is requested. So, as such, this is the reason why KNN is often referred to as a lazy learning algorithm.

So, this was all about the theoretical session. Now, let's move on to the coding part. So, for the practical implementation of the hands-on part, I'll be using the Iris dataset. So, this dataset consists of 150 observations. We have four features and one class label. The four features include the sepal length, the sepal width, petal length, and the petal width. Whereas the class label decides which flower belongs to which category. So, this was the description of the dataset which we are using.

Now, let's move on and see what are the step-by-step solutions to perform a KNN algorithm. So, first, we'll start by handling the data. What we have to do? We have to open the dataset from the CSV format and split the dataset into train and test parts. Next, we'll check the similarity, where we have to calculate the distance between two data instances. Once we calculate the distance, next, we'll look for the neighbor and select K neighbors which are having the least distance from a new point. Now, once we get our neighbor, then we'll generate a response from a set of data instances. So, this will decide whether the new point belongs to class A or class B. Finally, we'll create the accuracy function, and in the end, we'll tie it all together in the main function.

So, let's start with our code for implementing the KNN algorithm using Python. I'll be using Jupyter Notebook, Python 3.0 installed on it. Now, let's move on and see how the KNN algorithm can be implemented using Python. So, there's my Jupyter Notebook, which is a web-based interactive computing notebook environment with Python 3.0 installed on it. Let's launch. Yeah, it's launching. So, there's our Jupyter Notebook, and we'll be writing our Python codes on it. So, the first thing that we need to do is load our file. Our data is in CSV format without a header line. With any code, we can open the file with the `open` function and read the data line using the `reader` function in the CSV module. So, let's write a code to load our data file. Let's execute the run button. So, once you execute the run button, you can see the entire training dataset as the output.

Next, we need to split the data into a training dataset that KNN can use to make predictions and a test dataset that we can use to evaluate the accuracy of the model. So, we first need to convert the float measures that we loaded as strings into numbers that we can work with. Next, we need to split the dataset randomly into train and test. A ratio of 67 is to 33 for test is to train is a standard ratio which is used for this purpose. So, let's define a function as `load_dataset` that loads a CSV with a provided file name and splits it randomly into training and test datasets using the provided split ratio. So, this is our function `load_dataset` which is using `file_name`, `split_ratio`, `training_dataset`, and `testing_dataset` as its input. All right. So, let's execute the run button and check for any errors. So, it's executed with zero errors. Let's test this function. So, this is our training set, testing set, `load_dataset`. So, this is our function `load_dataset`, and inside that, we are passing our file `iris.data` with a split ratio of 0.66 and `training_dataset` and `test_dataset`. Let's see what our training dataset and test dataset it's dividing into. So, it's giving a count of `training_dataset` and `testing_dataset`. The total number of training dataset it has split into is 97, and the total number of test dataset we have is 53. So, the total number of training dataset we have here is 97, and the total number of test dataset we have here is 53. All right. Okay. So, our function `load_dataset` is performing well.

So, let's move ahead to step two, which is similarity. So, in order to make predictions, we need to calculate the similarity between any two given data instances. This is needed so that we can locate the K similar data instances in the training dataset and, in turn, make a prediction. Given that all four sepal measurements are numeric and have the same unit, we can directly use the Euclidean distance measure, which is nothing but the square root of the sum of squared differences between two arrays of numbers. Given that all the four flower measurements are numeric and have the same unit, we can directly use the Euclidean distance measure, which is nothing but the square root of the sum of square differences between two arrays of numbers. Additionally, we want to control which fields to include in the distance calculation. So, specifically, we only want to include the first four attributes. So, our approach will be to limit the Euclidean distance to a fixed length. All right.

So, let's define our Euclidean function. So, this is our Euclidean distance function which takes `instance1`, `instance2`, and `length` as parameters. `instance1` and `instance2` are the two points of which you want to calculate the Euclidean distance. Whereas this `length` denotes that how many attributes you want to include. Okay. So, there's our Euclidean function. Let's execute it. It's executing fine without any errors. Let's test the function. Suppose the data 1, or the first instance, consists of the data point as (2,2,2) and it belongs to class A, and data 2 consists of (4,4,4) and it belongs to class B. So, when we calculate the Euclidean distance of data 1 to data 2, and what we have to do, we have to consider only the first three features of them. All right. So, let's print the distance. As you can see here, the distance comes out to be 3.464. All right. So, this is nothing but the square root of (4-2)^2 + (4-2)^2 + (4-2)^2. So, this distance is nothing but the Euclidean distance, and it is calculated as the square root of (4-2)^2 + (4-2)^2 + (4-2)^2, that is nothing but 3 * (4-2)^2, that is 12, and the square root of 12 is nothing but 3.464. All right.

So, now that we have calculated the distance, now we need to look for K nearest neighbors. Now that we have a similarity measure, we can use it to collect the K similar instances for a given unseen instance. Well, this is a straightforward process of calculating the distance for all the instances and selecting a subset with the smallest distance value. And now, what we have to do, we have to select the smallest distance values. So, for that, we'll be defining a function as `get_neighbors`. So, for that, what we'll be doing, we'll be defining a function as `get_neighbors`. What it will do? It will return the K most similar neighbors from the training set for a given test instance. All right. So, this is how our `get_neighbors` function looks like. It takes `training_dataset`, `test_instance`, and `K` as its input. Uh, the K is nothing but the number of nearest neighbors you want to check for. All right. So, basically, what you'll be getting from this `get_neighbors` function is K different points having the least Euclidean distance from the test instance. All right. Let's execute it. So, the function executed without any errors. So, let's test our function. So, suppose the training dataset includes the data like (2,2,2) and it belongs to class A, and other data includes (4,4,4) and it belongs to class B, and our testing instance is (5,5,5). And now we have to predict whether this test instance belongs to class A or it belongs to class B. All right. For K = 1, we have to predict its nearest neighbor and predict whether this test instance it will belong to class A or will it belong to class B. All right. So, let's execute the run button. All right. So, on executing the run button, you can see that we have output as (4,4,4) and B. Our new instance (5,5,5) is closest to (4,4,4) which belongs to class B. All right.

Now, once you have located the most similar neighbor for a test instance, the next task is to predict a response based on those neighbors. So, how we can do that? Well, we can do this by allowing each neighbor to vote for their class attribute and take the majority vote as a prediction. Let's see how we can do that. So, we have a function as `get_response` which takes `neighbors` as the input. Well, this `neighbors` was nothing but the output of this `get_neighbor` function. The output of `get_neighbor` function will be fed to `get_response`. All right. Let's execute the run button. It's executed. Let's move ahead and test our function `get_response`. So, we have neighbors as (1,1,1) which belongs to class A, (2,2,2) it belongs to class A, (3,3,3) it belongs to class B. So, this response, what it will do, it will store the value of `get_response` by passing this `neighbor` value. All right. So, what we want to check is, we want to predict whether our test instance (5,5,5) it belongs to class A or class B when the neighbors are (1,1,1) A, (2,2,2) A, and (3,3,3) B. So, let's check our response.

Now that we have created all the different functions which are required for a KNN algorithm, the important main concern is how to evaluate the accuracy of the prediction. An easy way to evaluate the accuracy of the model is to calculate a ratio of the total correct predictions to all the predictions made. So, for this, I'll be defining a function as `get_accuracy` and inside that, I'll be passing my `test_dataset` and the `predictions`. `get_accuracy` function. Check it, executed without any error. Let's check it for a sample dataset. So, we have our test dataset as (1,1,1) which belongs to class A, (2,2) which again belongs to class A, (3,3,3) which belongs to class B. And my predictions is for the first test data, it predicted that it belongs to class A, which is true. For the next, it predicted that it belongs to class A, which is again true. And for the next, again it predicted that it belongs to class A, which is false in this case, 'cause the test data belongs to class B. All right. So, in total, we have two correct predictions out of three. All right. So, the ratio will be 2/3, which is nothing but 66.66%. So, our accuracy rate is 66.66%.

So, now that you have created all the functions that are required for the KNN algorithm, let's compile them into one single main function. All right. So, this is our main function, and we are using the Iris dataset with a split of 0.6767, and the value of K is three. Let's see what is the accuracy score of this. Check how accurate our model is. So, in the training dataset, we have 113 values, and in the test dataset, we have 37 values. These are the predicted and the actual values of the output. Okay. So, in total, we got an accuracy of 97.29%, which is really very good. All right.

Did Alibaba just do the impossible? Their latest AI model has outperformed both GPT-4 and DeepSeek in some key benchmarks. But how did they manage to do it? And what does this mean for the AI race? Stick around as we dive into the shocking details behind this breakthrough and what it means for the future of AI. Alibaba just dropped a bombshell during the Lunar New Year, a new AI model called Qwen 2.5 Max, and they say it outperforms OpenAI's GPT-4o, Meta's Llama, and even China's own rising star, DeepSeek. Is this the new benchmark in AI? Let's break it down. First off, what exactly is Qwen 2.5 Max? Developed by Alibaba Cloud, this model is being hyped as a major rival to GPT-4. According to their benchmarks, it crushes competitors in reasoning, coding, and multilingual tasks. So, let's look at the numbers. In Arena Hard, a benchmark for complex problem-solving, Qwen 2.5 Max scored 85.3%, beating GPT-4's 80.2% and DeepSeek V3's 77.5%. But here's the twist. It's not just about the raw power. Alibaba built this model for businesses. Think customer service bots that speak 10 languages or AI coders that debug Python faster than your engineering team. And unlike OpenAI's premium pricing, Alibaba's offering Qwen 2.5 Max at a fraction of the cost. But why drop this during the Lunar New Year when half of China's on vacation? Well, that's where the discussion begins. Meet DeepSeek, the 20-month-old startup that's been shaking up Silicon Valley. Three weeks ago, they dropped DeepSeek V3 and the R1 model. And the secret is insanely low cost. We are talking $0.14 million per token. That's like charging pennies for a Lamborghini. DeepSeek's open-source models triggered an AI price war in China. Alibaba reduced prices by up to 97% overnight. But DeepSeek founder Yang isn't sweating it. In a rare interview, he said, "We don't care about the price. AGI is our goal." AGI, that's Artificial General Intelligence. AI that can outthink humans. And here's the twist. DeepSeek isn't some corporate giant. They are a tiny team of grad students and researchers working out of Alibaba's hometown, Hangzhou. Meanwhile, Alibaba's got 200,000 employees. So, how does Qwen 2.5 Max actually stack up? Let's compare. First, let's talk about reasoning. Qwen takes the lead here. Next, when it comes to coding, Qwen continues to shine. And if multilingual support is what you are after, Qwen speaks Mandarin, English, Spanish, or whatever you name it. But DeepSeek V3 still holds the crown for affordability. And GPT-4, it's holding on to its reputation. But even OpenAI's Sam Altman admitted DeepSeek's progress is impressive. But Alibaba's timing is strategic. Releasing Qwen 2.5 Max during Lunar New Year when everyone's distracted is a power move. It's like dropping a diss track on Christmas Day. No one's looking, but everyone will hear it. So, here's why this matters. DeepSeek's R1 model wiped $1 trillion of US tech stocks in a day. Nvidia, Meta, Microsoft all dropped. Why? Because if a tiny Chinese startup can match GPT-4 at 1/100th the cost, investors wonder, are we overspending on AI? And China's giants aren't sitting still. ByteDance updated its AI model days after DeepSeek's launch. Tencent and Baidu are in the price-cutting frenzy. Meanwhile, Alibaba's betting big on Qwen to dominate enterprise AI. Think hospitals, banks, and mega-corporations. And the real question is, who's closer to AGI? DeepSeek's agile team or Alibaba's corporate powerhouse? Share your thoughts in the comments. So, is Qwen 2.5 Max the new AI champion? Maybe. But this isn't just about benchmarks. It's a glimpse into the future. A future where AI isn't just built by Silicon Valley giants, but by startups in Hangzhou and open-source communities worldwide. A game-changing development has taken the tech world by surprise. I think we should take the developments out of China very, very seriously. A free, open-source AI model emerged seamlessly out of nowhere. It not only matched but surpassed some of the most advanced systems on the market. What made this even more remarkable was its origin. It wasn't a new release from OpenAI, nor a breakthrough from Anthropic. It was a DeepSeek AI model developed in China, and its development left top AI researchers in the United States in amazement, especially when they learned about the staggering cost behind it. It's opened a lot of eyes of like, what is actually happening in AI in China. The training cost for DeepSeek version 3 was just $5.576 million. And in comparison, OpenAI spends a massive $5 billion annually. While Google's capital expenditures are projected to exceed $50 billion by 2024. Microsoft, on the other hand, invested over $13 billion just in OpenAI, and yet DeepSeek's model outperformed these highly funded AI models from leading American companies. And the contrast is truly mind-blowing to see the DeepSeek, um, um, new model. It's, it's super impressive in terms of both how they have really effectively done an open-source model that does what, uh, is this inference time compute, and it's super compute efficient. DeepSeek didn't stop at the success of its powerful open-source AI model. Instead, it quickly introduced R1, a next-generation reasoning model that has already surpassed the advanced OpenAI W1 model in several third-party benchmarks. This rapid innovation highlights DeepSeek's ability to surpass even the most well-funded United States AI giants, proving that agility and creativity can disrupt the established leaders in the race for AI dominance. As we dive deeper into this, let's hear from Martin Wishop, the Director of Bulgarian Institute for Computer Science, Artificial Intelligence, and Technology. He recently made some interesting statements about the AI industry, and they are shaking things up. So, he pointed out that a Chinese AI startup claimed to have developed its R1 LLM model with less than $6 million, while other companies are pouring in billions. That announcement alone caused Nvidia's stock to drop. And according to Martin, these models are built by strong researchers and engineers in the field, many of whom actively publish their work. But developing these AI models can be incredibly expensive. Just to give you an idea, running 2048 H800 GPUs could cost anywhere between $50 to $100 million. And he also mentioned that the company handling the data center is backed by a massive Chinese investment fund with far more GPUs than just those 2048 H800 units. As for the architecture behind DeepSeek R1 and V3 models, Martin explained that they use a Mixture of Experts (MoE) approach. Simply put, this means that at any given time, only a small percentage of the model is active, making it much more efficient in real-time use. This raises a lot of questions about cost-efficiency, AI development strategies, and how companies are competing in this space. Now, the question is, if DeepSeek's development is being reported to have cost only 5 to 6 million, how does this figure align with the extensive infrastructure, data center operations, and substantial backing from Chinese investment funds? Could there be more to the story that isn't being disclosed? Let us know in the comment section below. As far as the research indicates, DeepSeek V3 has been utilized as the base model for DeepSeek R1, and this progression highlights DeepSeek's strategic approach to building on its existing architecture while pushing the boundaries of AI capabilities. DeepSeek R1 distinguishes itself by relying entirely on reinforcement learning fine-tuning, a focused and efficient method that contrasts sharply with OpenAI's GPT infrastructure. OpenAI's GPT (Generative Pre-trained Transformer) framework employs a combination of supervised learning, unsupervised learning, and reinforcement learning to train its models. While this multifaceted approach has proven effective, it also requires significant computational resources and time. In contrast, DeepSeek R1's exclusive use of reinforcement learning fine-tuning demonstrates a more streamlined and targeted methodology, which not only reduces cost but also enhances performance in specific tasks. These differences in training strategies highlight DeepSeek's remarkable ability to innovate efficiently. And by building on its foundational V3 model, it developed R1, a reasoning model that has already surpassed OpenAI's advanced systems in some key benchmarks. DeepSeek's focus on reinforcement learning has allowed it to carve out a unique position, directly challenging the dominance of United States AI giants. This approach demonstrates that with strategic, resource-conscious innovation, groundbreaking results are not only possible, but they are already happening. So, what do you think? Does DeepSeek's open-source approach give it a long-term advantage? Or will OpenAI's heavy investment in research and proprietary models keep it ahead? Share your thoughts in the comments. Did a Chinese AI model just shake up the entire US market? Let me tell you what happened. On January 27, 2025, the stock market took a serious hit. Tech stocks dropped hard, and the biggest shock was that Nvidia, a prominent player in the AI hardware sector, saw its stock crash by 17%, and that's a $590 billion loss in market value. And it was because of an AI model called DeepSeek R1. Yes, you heard that right. A Chinese AI model just sent shockwaves through the industry, raising big concerns about China's growing AI dominance and what it means for companies like OpenAI, Google, and even Nvidia. So, why is DeepSeek R1 such a big deal? Well, it's not just another AI model. It's a game-changer. And with models like DeepSeek R1, DeepSeek V2, and DeepSeek Coder, it's going head-to-head with the top players like OpenAI and Google, offering powerful AI at a fraction of the cost. Here, I have added a screenshot for your reference, and this table compares the large language models based on their accuracy and calibration error. And among the models mentioned, DeepSeek R1 has the highest accuracy with 9.4%, outperforming 01 with 9.1%, Gemini Thinking with 6.2%, and other models such as GPT-4o with 3.3% and Group 2 with 3.8%. Furthermore, DeepSeek R1 has the lowest calibration error with 81.8%, indicating improved confidence calibration over other models with errors greater than 88%. This demonstrates that DeepSeek R1 not only produces the most accurate results but also has a higher forecast reliability. And this benchmark graph will show the DeepSeek R1's exceptional performance across a variety of evaluation tasks, solidifying its position as a top-tier LLM. Notably, DeepSeek R1 achieves the highest scores in AIM 2024 with 79.8%, Codeforces with 96.3%, Math 500 with 97.3%, and MMLU with 90.8%, indicating superior reasoning, problem-solving, and coding skills when compared to OpenAI's 01 models and DeepSeek V3. DeepSeek R1 consistently outperforms or equals top models, particularly in domains requiring precise logical reasoning and mathematical skills. Furthermore, its highest Viven verified score is 49.2%, demonstrating its suitability for software engineering applications, and this finding supports DeepSeek R1's advancements in AI research, establishing it as a formidable competitor in the LLM space.

First, let's talk about DeepSeek R10 and its successor, DeepSeek R1. So, let's break it down. In reinforcement learning, there are two main components: the agent and the environment. The agent interacts with the environment, and based on its actions, it receives rewards or penalties. The goal of the agent is to maximize these rewards by learning from its mistakes and improving over time. Now, let's talk about the DeepSeek R10. This model was a pioneering attempt to use reinforcement learning without supervised fine-tuning. And the idea was to let the model learn entirely through interaction with its environment without any pre-labeled data. However, this approach had some challenges. DeepSeek R10 faced two major issues. First is the poor readability. The model's outputs were often hard to understand. And next, the language mixing with Chinese. The model sometimes mixed languages, especially Chinese, which affected its performance in English tasks. And to address these issues, the team introduced DeepSeek R1. This new model not only solved the problems of readability and language mixing but also achieved remarkable performance. In fact, DeepSeek R1 matched the accuracy of OpenAI's GPT-01 model, especially the OpenAI 01217 model on reasoning tasks. That's a huge milestone. But that's not all. DeepSeek R1 is also 24 to 28 times cheaper to train compared to other state-of-the-art models. And this makes it not only highly effective but also cost-efficient, opening up new possibilities for research and applications. So, to recap, here we have the DeepSeek R10 was an ambitious attempt at reinforcement learning without supervised fine-tuning, but it faced challenges with readability and language mixing. DeepSeek R1 addressed these issues, achieving top-tier accuracy and being significantly more cost-effective. Now, as far as GPT is concerned, ChatGPT combines unsupervised learning, supervised fine-tuning, and RLHF, making it more aligned for text-based reasoning and safe AI interactions.

Now, we will look at the model comparison. DeepSeek and GPT are both pushing the boundaries of what AI can do, but they take very different approaches. So, let's break it down. We asked the GPT-01 model and DeepSeek R1 model to generate a Python code where a ball bounces inside a rotating triangle. Sounds cool, right? Well, let's check out the result. First up, here's what GPT-01 came up with. It works, but the physics seems a bit off, and the movement isn't as smooth as you would expect. Not bad, but it's not quite there yet. Now, let's look at what DeepSeek R1 generated. Wow, this one looks way better, right? The ball's movement feels more natural, and the rotation of the triangle is much smoother. The overall gameplay experience is just more polished. So, if we compare the two, DeepSeek R1 definitely outperformed GPT-01 in this challenge. Of course, both models are impressive in their own ways, but when it comes to designing this specific game, DeepSeek R1 takes the win.

Now, we will distinguish the differences in detail. So, first, let's talk about the architecture. DeepSeek uses a Mixture of Experts design. Think of it like a team of specialists. Only the relevant experts are activated for each task. So, for example, DeepSeek V3 has 671 billion parameters, but only 37 billion are activated per token, making it super efficient. And on the other hand, GPT models use a dense transformer architecture where all the parameters are active at once. GPT-4, for instance, has 175 billion parameters, all working simultaneously, and this makes GPT powerful but also computationally expensive.

Now, let's talk about cost and efficiency. DeepSeek is a game-changer here. It was developed on a budget of just $5.5 million due to its efficient MoE design, and that's a fraction of what other models cost. GPT models like GPT-4 require massive computational resources. Training GPT-4 cost over a hundred billion dollars, making it a heavyweight in terms of both performance and expense. And when it comes to performance, both models shine in different areas. DeepSeek is a powerhouse in tasks like coding, translation, and solving complex math problems. And in fact, DeepSeek R1 has been shown to match the performance of advanced systems from OpenAI and Google despite its smaller budget. GPT models like GPT-4 are known for their natural language understanding, creative writing, and complex reasoning, and they are incredibly versatile and can handle a wide range of tasks with ease.

Next, accessibility is another key difference. DeepSeek is open-source, meaning its code is available to the public, and this promotes transparency, collaboration, and innovation within the AI community. GPT models, on the other hand, are primarily proprietary. While OpenAI has released some tools and models, many advanced versions are restricted and accessed through APIs.

Finally, let's talk about ethics and censorship. DeepSeek implements strict content moderation, especially for politically sensitive topics. This ensures compliance with regulatory standards but can sometimes limit its responses. GPT models also have moderation mechanisms to prevent harmful outputs, but they strive to balance open access to information with ethical guidelines.

Now, we will install the DeepSeek R1 and run a short demo on it. So, let's see how the DeepSeek R1 model can be installed. First, let's open up our browser and head over to the Ola.com. And once you're there, you will see a download button. So, go ahead and click on that. Now, select "Download for Windows" to start the download. And keep in mind, it's a pretty big file, so it might take a minute to download. So, let's give it some time. Once the download finishes, go to the download folder and find the installation file. Now, here, double-click on the file to open the installation window, and you will see an "Install" button. So, simply click on that, and Ola will start installing on your system. So, once Ola is installed, let's head back to the Ola website. Now, click on the "Models" tab at the top of the page, and here you will see a list of available models. And for this video, we are going to use the DeepSeek R1 model. So, we will select the 1.5B model, but if you want, you can also choose the latest 7B model too. Now, once you have selected the model, you will find the installation command here. So, go ahead and copy that command. Now, open your terminal on your Windows system and paste the command we copied earlier. So, this is the command, and this command will start pulling the model. So, depending on your internet speed, this might take a little time. So, be patient, and that's it. Once the process is done, your DeepSeek R1 model will be all set up and ready to use on your system. All right.

Now, let's try out some commands here. So, let's say "hello." Okay, we got some response here. Now, let's ask it to tell something about itself. All right, so it responded saying that "I am DeepSeek R1, an AI assistant." Next, let's ask it to design a Python code to list all the files in a directory. So, as you can see, it has provided the code we asked for. Great, right?

So, what does all of this mean for the future of AI? Well, AI is becoming more accessible. Like, for years, AI development was mostly controlled by big companies with massive budgets. But now, with open-source models like DeepSeek R1, anyone, whether you have a small startup, or you are an independent developer, or a student, you can build AI solutions without paying huge API fees. This means faster adoption and innovation worldwide. Next, the global AI race is heating up. The AI race between the United States and China is getting even more competitive. While the United States tries to limit AI chip exports, China is finding ways to keep up. No matter which side you support, one thing is clear: AI is evolving rapidly, and staying informed is more important than ever. Next, AI is becoming more sustainable. Training AI models consumes massive amounts of energy, but with advancements in optimization, we are seeing a shift towards more efficient and eco-friendly AI. This means lower CO2 emissions and a reduced environmental impact, something that was once a major concern in AI development. Next, career opportunities are growing. And if you're a developer, AI engineer, or a data scientist, this is your moment. Companies will need skilled professionals to build and deploy AI solutions at a faster pace than ever. So, if you have been thinking about getting into AI, now is the time to start. Well, DeepSeek is making big moves, but can it really compete with OpenAI in the long run? Which one do you think will dominate the future of AI? So, let me know your thoughts in the comments below.

So, let's begin our deep learning interview question and answer session and understand what are the typical questions which are being asked in a deep learning interview. So, the first and foremost question what any deep learning interviewer asks is the basic understanding or the relationship between machine learning, artificial intelligence, and deep learning. So, basically, artificial intelligence is a technique which enables machines to mimic human behavior, and machine learning is a subset of artificial intelligence techniques which uses statistical methods to enable machines to improve with experience. Now, deep learning, on the other hand, is a subset of machine learning which makes the computational multi-layer neural network feasible. It uses neural networks to simulate human-like decision-making.

Now, coming to the second question: Do you think deep learning is better than machine learning, and if so, why? Now, though machine learning algorithms, the traditional machine learning algorithms, solve a lot of our cases, but they are not very useful while working with high-dimensional data. Now, that is where we have a large number of inputs and outputs. For example, in the case of handwriting recognition, we have a large amount of inputs where we have different types of input associated with different types of handwriting. Now, another major challenge is to tell the computer what all features it should look for that will play an important role in predicting the outcome, as well as to achieve better accuracy while doing so. So, these are some of the few shortcomings what machine learning has, and deep learning overcomes all of these shortcomings.

Now, coming to our third question, which is: What is a perceptron and how does it work? Now, actually, our brain has subconsciously trained itself to do a lot of things over the years. Now, the question comes, how does deep learning mimic the functionality of the brain? Well, deep learning uses the concept of an artificial neuron that functions in a similar manner as the biological neuron present in our brain. Therefore, we can say that deep learning is a subfield of machine learning concerned with algorithms inspired by the structure and function of the brain, called artificial neural networks. Now, if you focus on the structure of a biological neuron, it has dendrites which are used to receive inputs. Now, these inputs are summed in the cell body, and using the axon, it is passed on to the next biological neuron. Now, similarly, a perceptron receives multiple inputs, applies various transformations and functions, and provides an output. Now, a perceptron is a linear model used for binary classification. It models a neuron which has a set of inputs, each of which gives a specific weight. Now, the neuron computes some function on these weighted inputs and then finally provides the output. As we know that our brain consists of multiple connected neurons called the neural network, we can also have a network of artificial neurons called the perceptron to form a deep neural network.

Now, coming to the next question: What is the role of weights and biases? Now, for a perceptron, there can be one or more input called bias. While the weights determine the slope of the classifier line, the bias allows us to shift the line towards the left or right. And normally, bias is treated as another weighted input with the input value X. In our case, if you have a look at a typical perceptron, what it receives is a set of inputs. Now, these inputs are not just inputs which it gathers. So, weights are an additional input which it takes, and according to that, it computes and provides an output.

Now, which brings us to the next question, which is: What exactly are activation functions? So, an activation function translates the inputs into outputs and it uses a threshold to produce an output. So, the activation function decides whether a neuron should be activated or not by calculating the weighted sum and further adding the bias to it. And the purpose of the activation function is to introduce a nonlinearity into the output of a neuron. There can be many activation functions like linear or identity. We have the binary step, we have sigmoid, we have tanh, we have ReLU, and softmax. These are a lot of activation functions which are being heavily used in the deep learning industry. So, one should actually know about all of these things.

Now, talking about perceptron, our next question what an interviewer might ask is: Explain the learning of a perceptron. So, basically, a perceptron has four steps of learning. So, the first step is initializing the weights and threshold. So, just now, as I mentioned, initializing the weights and the threshold so to the perceptron so that it can activate a neuron by calculating the weighted sum and further adding the bias in and all. This is the first step. And the second step is providing the input and calculating the output using the activation functions. And according to that, what we do is, the third step involves updating the weights. Now, once a particular perceptron learns something, it has to update the weights so that it can learn much more things in a new manner. And the next step what comes is, repeat step number two and three, which is, provide the input and calculate the output, and then update the weights accordingly. Now, if you have a look at the equation here, we have w_j(t+1) = w_j(t) + eta * (y - y_hat) * x_j. The w_j(t+1) is the updated weight, whereas w_j(t) is the old weight, y is the desired output, y_hat is the actual output, and x_j is the input. So, this is the equation of the learning of a perceptron.

Now, the next question is: What is the significance of a cost or a loss function? So, a cost function is a measure of accuracy of the neural network with respect to a given training sample and expected output. It provides the performance of a neural network as a whole. And in deep learning, the goal is to minimize the cost function. So, for that, we use the concept of gradient descent.

Now, which brings us to the next question, which is: What exactly is gradient descent and what are its various types? So, gradient descent is an optimization algorithm which is used to minimize some function by iteratively moving in the direction of the steepest descent, as defined by the negative of the gradient. Now, think of it as a bowl in which you start from any particular point, and the goal is to reach the bottom of the bowl, which is the gradient descent. So, there are uh, three types of gradient descent, which are the stochastic, batch, and the mini-batch. So, stochastic gradient descent uses only a single training example to calculate the gradient and update the parameters accordingly. Whereas the batch gradient descent calculates the gradients for the whole dataset and performs just one update at each iteration. Now, mini-batch gradient descent is a variation of the stochastic gradient descent where instead of a single training example, a mini-batch of samples are used, and it is one of the most popular optimization algorithms.

Now, if we talk about mini-batch gradient descent, one might ask, is what are the benefits of the mini-batch gradient descent or how is it useful than the others? Now, the mini-batch gradient descent is more efficient when compared to the stochastic gradient descent, and the generalization is done by finding the flat minima, which allows to help approximate the gradient of the entire training set, which helps us to avoid the local minima. Now, this is why mini-batch gradient descent is considered or is preferred over the regular gradient descent algorithm, which is the stochastic gradient descent.

Now, one might ask, what are the steps for using a gradient descent algorithm? So, first of all, what you need to do is initialize some random weights and biases. And after that, you need to do is pass an input through the network and get values from the output layer. Next, what you're going to do is calculate the error between the actual value and the predicted value. Now, this can be done in a number of ways. Now, the next step involves is to go to each neuron which contributes to the error and change its respective values to reduce the error, which is basically our goal is to reduce the cost of any particular function or any particular model. So, after that, what you do is reiterate until you find the best weights of the network and you find the lowest cost of the particular network.

So, one might ask you to write any gradient descent program or write the pseudo-code of any gradient descent program. So, what you need to do, first of all, what we do is define the parameters, which are the weights, the hidden weights, the weight output, the bias hidden, and the bias output. We define a function `std` with arguments as `cost`, the `parameters` what we have discussed, and the `learning_rate`. Now, what we do is then we then define the gradients of our parameters with respect to the cost function. So, here we use the Theano library to find the gradients, and we import Theano as T, and finally iterate through all the parameters to find out the updates for all the possible parameters. So, you can see that we use vanilla gradient descent here, and as you can see, it returns the updates, and what we do is update the parameters and the cost in this particular equation. The ultimate goal of any gradient descent algorithm is to minimize the cost.

Now, talking about perceptron, what are the shortcomings of a single-layer perceptron? So, well, there are two major problems. Now, first of all, is that the single-layer perceptron cannot classify non-linearly separable data points. And the second point is that the complex problems that involve a lot of parameters cannot be solved by a single-layer perceptron. Now, consider an example here and the complexity which arises when the parameters are involved to take a decision by a marketing team. So, first of all, we have the categories which are the email, direct, paid, referral program, or the organic. And inside these categories, we have subcategories which are the Google, Facebook, LinkedIn, Twitter, we have Instagram now. And inside that, we have the type of subcategory which are the search ads, remarketing ads, interested ads, lookalike ads. And again, if we do a subdivision, we have the parameters to consider, which are the customer acquisition cost, we have the money spent, and the click rate or the lead generated, the customer generated, and the time taken to become a customer. So, one neuron cannot take in so many inputs, and that is why more than one neuron would be used to solve this problem.

Now, which brings us to the question: What is a multi-layer perceptron? So, a multi-layer perceptron, or MLP, is a class of feedforward artificial neural network, and it is composed of more than one perceptron. They are composed of an input layer to receive the signal, an output layer that makes a decision or the prediction about the input, and in between these two, an arbitrary number of hidden layers that are the true computational engine of any multi-layer perceptron.

Now, one might ask, what are the different parts of any multi-layer perceptron or a neural network? So, first of all, what we have are input nodes. So, uh, the input nodes provide information from the outside world to the network and are together referred to as the input layer. No computation is performed in any of the input nodes.

Nodes. They just pass the information to the hidden layers. Now, hidden nodes have no direct connection with the outside world, hence the name "hidden." And what they do is they perform computation and transfer the information from the input nodes to the output nodes. Now, a collection of hidden nodes forms the hidden layer, and while a network will only have a single input layer and a single output layer, it can have zero to n number of hidden layers. A multi-layer perceptron has more than one hidden layer.

Now, if we talk about output nodes, the output nodes are collectively referred to as the output layer and are responsible for the computation and transferring information from the network to the outside world. Hence, they are also responsible for the prediction.

Now, coming to our next question: What exactly is data normalization and why do we need it? Now, data normalization is a very important pre-processing step, which is to normalize the data. The data should not be either left-skewed or right-skewed; it should be normal. It is used to rescale the values to fit in a specific range to assure better convergence during backpropagation. In general, it boils down to subtracting the mean of each data point and dividing by its standard deviation so that we get a normally distributed data, and it makes computation easy in terms of backpropagation in the case of neural networks. So, this is a very important part of any deep neural network.

Now, talking about deep neural networks, or neural networks in general. Coming to our next question, which is: Now, what is better, the deep networks or the shallow ones, and why? Now, both networks, be it shallow or deep, are capable of approximating any function. What matters is how precise that network is in terms of getting the result. Now, a shallow network works with only a few features, as it cannot extract more. But a deep network goes deep by computing efficiently and working on more features or parameters. Now, deeper networks are able to create deep representations at every layer. The network learns a new, more abstract representation of the input, and hence deep neural networks are better than the shallow ones.

So, what exactly is weight initialization in a neural network? Now, as we saw, we had weight initialization in perceptron. So, weight initialization is one of the very important steps. A bad weight initialization can prevent a network from learning, but good weight initialization can help it in giving quicker convergence and a better overall error. Now, biases can be generally initialized to zero. The rule for setting the weights is to be close to zero without being too small, because every time the weight is being multiplied to the inputs, the result gets smaller and smaller.

Now, talking about neural networks, what is the difference between a feed-forward and a backpropagation neural network? Now, a feed-forward neural network is a type of neural network architecture where the connections are fed forward, that is, they do not form cycles. The term "feed-forward" is also used when you input something at the input layer, and it travels from the input to the hidden and from the hidden to the output layer. The values are fed forward.

Now, backpropagation is a training algorithm which consists of two steps majorly. The first one is feed-forwarding the values, and the second one is to calculate the error and propagate it back to the earlier layers. So, to be precise, forward propagation is a part of the backpropagation algorithm, but it comes before the backpropagation.

So, one might ask the question, which is one of the most important questions: What are the hyperparameters in a neural network, and name a few of these hyperparameters? So, hyperparameters are the variables which determine the network structure, that is, for example, the number of hidden units and or the hidden layers, and the variables which determine how the network is trained, for example, the learning rate.

Now, there are two types of hyperparameters usually. One are the network parameters which are associated with the network. In that case, we have the number of layers, we have the network weight initialization, we have the activation function. And in the training parameters, we have the learning rate, we have momentum, number of epochs, we have the batch size, and much more.

Now, a lot of hyperparameters also differ when we work along with different types of neural networks. So, as in CNN, we get extra parameters to work on when considering CNN, which are the convolutional neural networks, and sometimes we have to deal with fewer hyperparameters. It all depends upon the type of neural network which you are using.

So, uh, which brings us to the next question: Explain the different hyperparameters related to networking and training. So, in training, we have, first of all, we have the number of hidden layers. So, hidden layers are the layers between the input and the output layers, as we just discussed, and many hidden units within a layer with regularization technique can increase the accuracy, as a smaller number of units may cause underfitting.

Now, another important aspect is network weight initialization. So, ideally, it may be better to use different weight initialization schemes according to the activation function used on each layer. Mostly, uniform distribution is used or the normal distribution.

Now, if we talk about activation function, so they are also used to introduce nonlinearity to the models. They are also used to introduce nonlinearity to the models, which allows deep learning models to learn nonlinear prediction boundaries. Now, generally, the rectifier activation function, or the ReLU, is the most popular.

Now, if we talk about the training parameters. So, these were the network parameters which have to be initialized to a deep neural network before the training begins. And just before the training, we have the training parameters, which are the learning rate. So, the learning rate defines how quickly a network updates its parameters. Low learning rate slows down the learning process but converges smoothly. A larger learning rate speeds up the learning but may not converge as smoothly as a low learning rate. Usually, a decaying learning rate is preferred so that we get the best of both worlds and we get the best expected output.

Now, another hyperparameter is momentum. So, momentum helps us to know the direction of the next step with the knowledge of the previous step. Now, it helps to prevent oscillation, and a typical choice of momentum is between 0.5 to 0.9.

Now, if we talk about the number of epochs. So, epoch is basically iterations. So, number of epochs is the number of times the whole training data is shown to the network while training. So, increase the number of epochs until the validation accuracy starts decreasing, your training accuracy is increasing. So, that results in sometimes overfitting.

And if we talk about the batch size, so mini-batch size is the number of subsamples given to the network after which parameter updates happen. So, a good default for batch size might be 32 or 16, 64. It depends upon the size of, you know, the data you have. It can be any arbitrary number, but it's always better to have it in the power of two, right?

So, while we were talking about overfitting, which brings us to our next question, which is: What exactly is dropout? So, dropout is a regularization technique to avoid overfitting, which is to increase the validation accuracy, thus increasing the generalization power. Now, generally, use a small dropout value of 20% to 50% of the neurons, with 20% providing a good starting point. A probability too low has minimal effect, and a value too high results in underlearning by the network. So, first of all, what you need to do is use a large network, and you are likely to get better performance when the dropout is used on a larger network, giving the model more of an opportunity to learn independent representations.

Now, our next question is: In a neural network, you notice that the loss does not decrease in the few starting epochs. So, what could be the possible reason for this to happen? Now, the correct answer is: The reason for this could be the learning rate is low, first of all, or it might be the regularization parameter is high, or it can be it is stuck at local minima. So, it might take certain iterations to go out of that local minima and finally reach the lowest point. So, it might happen in some cases that it is stuck at local minima. So, another approach to that sort of problem must be initiated at that particular point of time.

Now, talking about deep learning, one might ask to name a few deep learning frameworks which are being used in the industry. So, first of all, the foremost and the most amazing deep learning library is TensorFlow. Followed by, we have Caffe. We have the Microsoft Cognitive Toolkit, which is the CNTK. We have Torch or PyTorch, which is giving a good battle or it's standing out from the crowd, and people are sometimes preferring PyTorch over TensorFlow. Now, MXNet is another deep learning framework. We have Chainer, and we have Keras. Now, Keras, as you know, can be integrated with Theano as well as TensorFlow, and Keras has been considered one of the best or the simplest deep learning frameworks when it comes to deep learning.

Now, one might ask, what exactly are tensors? So, tensors are nothing but a de facto for representing the data in deep learning. What I meant to say is that tensors are just multi-dimensional arrays that allow you to represent data having higher dimensions. In general deep learning, you deal with high-dimensional datasets where dimension refers to the different features present in the dataset. So, what you need is a multi-dimensional sort of array or a data structure, what you could say. So, that's what exactly a tensor is, and in fact, the name TensorFlow has been derived from the operations which the neural network performs on tensors. So, it's literally a flow of tensors.

Now, talking about TensorFlow, one might ask, since it's the most popular deep learning framework and companies prefer people having the knowledge of TensorFlow and been working on it, so what are the few advantages of TensorFlow? So, first of all, it has platform flexibility. It is easily trainable on CPU as well as GPU for distributed computing. Now, TensorFlow has auto-differentiation capabilities, and it has advanced support for threads, asynchronous computation, and it is a customizable and open-source framework. And most importantly, if we talk about the latest TensorFlow 2.0, which has just been released, so those come up with a lot of interesting features, and it has adopted Keras as its high-level API fully, so that the coding aspect of it is much simplified, and eager execution is now by default, so that you do not have to write loads and loads of lines of code. And if you want to know more about TensorFlow 2.0 and why it's the best learning framework right now, just go ahead and check our TensorFlow 2.0 video. I'll leave the link in the description box below. Go check it out, guys, and understand how exactly it is better from the previous version and why it is the best deep learning framework right now.

Now, talking about computational graphs, one might ask what exactly they are. So, well, a computational graph is a series of TensorFlow operations arranged as nodes in the graph. Now, each node takes zero or more tensors as input and produces a tensor as output. Now, basically, one can think of a computational graph as an alternative way of conceptualizing mathematical calculations that take place in a TensorFlow program. Now, the operations assigned to the different nodes of a computational graph can be performed in parallel, thus providing better performance in terms of computation.

So, one might ask, what exactly is a Convolutional Neural Network? Now, a Convolutional Neural Network, or CNN, or ConNet, is a class of deep learning neural networks which is most commonly applied to analyzing visual imagery. So, CNNs use a variation of the multi-layer perceptron designed to require minimal processing. Now, one might ask the next question: If you are going for an interview which requires you to work with a lot of images or videos, so in that case, CNNs are very much used. So, having a good knowledge of CNN is always better in that case.

So, the next question what we have here is: What are the various layers of CNN? Now, there are four layered concepts everyone should understand in Convolutional Neural Networks: are first, the convolutional layer, the second is the ReLU layer, and finally, we have the pooling layer, and finally, we end up with the full connectedness or the fully connected layer.

Now, if we talk about CNN, we have to talk about RNN also. So, one might ask, what exactly is RNN? So, RNN, or Recurrent Networks, are a type of artificial neural network which are designed to recognize patterns in the sequence of data, such as text, genomes, handwriting, the spoken word, numerical time series data from sensors, the stock markets, and the government agencies. So, recurrent neural networks use the backpropagation algorithm for training, but it is applied for every timestamp. It is commonly known as backpropagation through time, which is BPTT.

Now, our next question is: What are some issues faced while training an RNN? So, recurrent neural networks use the backpropagation algorithm, as I just mentioned, for training, but it is applied for every timestamp, and there are some issues with backpropagation, such as vanishing gradient or the exploding gradient, where the gradient vanishes or it is too much to handle.

Which brings us to the next set of questions. The first of which is: What exactly is a vanishing gradient and how is it harmful? Now, when we do backpropagation, that is, move backward in the network and calculate gradients of loss, which is the error with respect to the weights, the gradients tend to get smaller and smaller as we keep on moving backward in the network. Now, this means that the neurons in the earlier layers learn very slowly as compared to the neurons in the later layers in the hierarchy. Now, the earlier layers in the networks are the slowest to train.

Now, how is this harmful? So, earlier layers in the neural networks are important because they are responsible to learn and detect the simple patterns and are actually the building blocks of our neural network. Obviously, if they give improper and inaccurate results, then how can we expect the next layer and the complete network to perform nicely and produce the accurate result? So, the training process takes too long, and the prediction accuracy of the model will decrease.

Now, another question here arises is: What exactly is then exploding gradient descent? Now, this is just the opposite of vanishing gradient descent. So, exploding gradients are a problem when large error gradients accumulate and result in very large updates to the neural network model weights during training. So, the gradients are used during the training to update the network weights. But typically, when this process works best is when these weights are small and controlled. When the magnitudes of the gradients accumulate, an unstable network is likely to occur. Now, which causes poor prediction and results, or even a model that reports nothing useful whatsoever. So, vanishing gradient and the exploding gradient are two problems which occur while the backpropagation happens in a recurrent neural network.

So, our next question is: What are LSTMs? So, Long Short-Term Memory, which are the LSTMs, is an artificial recurrent neural network architecture used in the field of deep learning. And unlike standard feed-forward neural networks, the LSTM has feedback connections that make it a general-purpose computer. Now, it can not only process single data points but also the entire sequence of data. They are a special kind of RNN, or the recurrent neural network, which are capable of learning long-term dependencies.

Now, one might ask, what are capsules in a capsule neural network? So, capsules are vector, or what we can say, an element with a size and a direction specifying the features of the object and its likelihood. Now, these features can be any of the instantiation parameters like the pose, we have the position, size, orientation, deformation, velocity, the albedo, which is the light reflection, hue, texture, and much more. A capsule can also specify its attributes like angle and size. So, it can represent with the same generic information.

Now, just like a neural network has layers of neurons, a capsule network can have layers of capsules. So, there could be higher capsules representing the group of objects or the capsules below them. Now, this helps in getting deeper knowledge of a particular object or a particular dataset and having the knowledge from different aspects or different angles.

So, the next question arises is: Explain autoencoders and its uses. So, an autoencoder neural network is an unsupervised machine learning algorithm that applies backpropagation setting the target values to be equal to the inputs. So, autoencoders are used to reduce the size of our inputs into smaller representations, and if anyone needs the original data, they can reconstruct it from the compressed data.

Now, one might ask the question: How does an autoencoder differ from PCA? So, an autoencoder can learn from nonlinear transformations with a nonlinear activation function and multiple layers. It does not have to learn dense layers. It can use convolutional layers to learn, which is better for video, image, and series data. It is more efficient to learn several layers with an autoencoder rather than learn one huge transformation with PCA. An autoencoder provides a representation of each layer as the output and can take the use of pre-trained layers from other models to apply transfer learning to enhance the encoder or the decoder. So, these are few of the reasons why autoencoders are better than PCA, as we know both of them perform the same task, which is mostly dimensionality reduction.

Now, give some real-life examples where autoencoders can be applied. So, first of all, we talk about dimensionality reduction. Or the first thing that should pop up in your mind is dimensionality reduction. So, the reconstructed image is the same as our input image but with reduced dimensions. Now, it helps in providing a similar image with reduced pixel values, and it can be used in various areas where we have limited storage or we have limited processing power. So, when there is a high input or an image or a data with high dimensions or which has higher values, pixel values, it can compress and provide the same image with a lower pixel value, right? Or colors are used for converting any black and white picture into a colored image. Believe it or not, and depending on what is in the picture, it is possible to tell what the color should be.

Now, feature variation. If we talk about feature variation, it attracts only the required features of an image and generates the output by removing any unnecessary noise or unnecessary interruption. And if we talk about denoising images, the input seen by an autoencoder is not the raw input but a stochastically corrupted version. A denoising autoencoder is thus trained to reconstruct the original input from the noisy version.

Now, talking about autoencoders, one might ask about the different layers of the autoencoders. So, basically, an autoencoder consists of three layers, which is the encoder, we have the code, and the decoder. Which brings us to the next question: Explain the architecture of an autoencoder. If you talk about the three layers, which are encoder, code, and decoder. So, if we talk about the encoder, this part of the network compresses the input into a latent space representation. Now, the encoder layer encodes the input images as a compressed representation in a reduced dimension, and the compressed image is the distorted version of the original image.

Now, coming to the middle part, which is the code. So, this part of the network represents the compressed input which is fed to the decoder. It is basically the channel. And if you talk about the decoder, this layer decodes the encoded image back into the original dimension. And a decoded image is a lossy reconstruction of the original image, and it's reconstructed from the latent space representation.

Now, one might ask, what exactly is a bottleneck in an autoencoder and why is it used? Now, the layer between the encoder and the decoder, that is the code, is also known as the bottleneck. So, this is a well-designed approach to decide which aspects of the observed data are relevant information and what aspects can be discarded. It does this by balancing two criteria: the first, the compactness of the representation measured as the compressibility, and second, it retains some behaviorally relevant variables from the input.

Now, one might ask, are there any variations of autoencoders? Surely, there are. So, there are conventional autoencoders, we have sparse autoencoders, we have deep autoencoders, we have contractive autoencoders. All of these autoencoders have a different structure or the different code layer. If you talk about the convolutional autoencoder, we have the convolutional CNN algorithm sort of structure in that particular autoencoder with the encoder on one side. We have the convolution layers, the ReLU layer, the pooling layer inside it, and then finally, we have the decoding layer.

So, another question what might pop into the interviewer's mind is: What are deep autoencoders? So, the extension of simple autoencoders is deep autoencoders. The first layer of the deep autoencoders is used for first-order features in the raw input. Now, the second layer is used for second-order features corresponding to the patterns in the appearance of the first-order features. So, the deeper layers of the deep autoencoders tend to learn even higher-order features. So, a deep autoencoder is composed of two symmetrical deep belief networks: first, four or five shallow layers representing the encoding half of the net, and the second set of four or five layers that make up the decoding half. Interesting, right?

So, another important topic in deep learning are the Restricted Boltzmann Machines. So, one might ask, what exactly is an RBM or Restricted Boltzmann Machine? So, RBM is an undirected graphical model that plays a major role in deep learning frameworks in recent times, and it is an algorithm which is used for dimensionality reduction. Not only that, it is used for classification, regression, collaborative filtering, feature learning, and topic modeling.

So, when we talk about RBM being useful for dimensionality reduction, another question might arise is: How does RBM differ from autoencoders? So, an autoencoder is a simple three-layer neural network where output units are directly connected back to the input units. Typically, the number of hidden units is much less than the number of visible ones, and the task of training is to minimize an error or the reconstruction, that is, find the most efficient compact representation for the input data.

So, RBMs share a similar idea, but it uses stochastic units with a particular distribution instead of deterministic distribution. The task of our training is to find out how these two sets of variables are actually connected to each other. One aspect that distinguishes RBM from autoencoders is that it has two biases. The hidden bias helps the RBM produce the activations on the forward pass, while the visible layer biases help the RBM learn the reconstruction on the backward pass.

Now, this brings us to the final question of our deep learning interview: What are some limitations of deep learning? I bet you weren't thinking of this one, but there are some limitations. So, deep learning usually requires large amounts of training data, and deep neural networks are easily fooled. Now, the success of deep learning is purely empirical. Deep learning algorithms have been criticized as uninterpretable black boxes because one important thing about deep learning is that you do not specify what you are looking for, right? The algorithm learns on its own. So, that is one of the shortcomings of deep learning, and deep learning thus far has not been well integrated with prior knowledge.

So, a lot of people still don't feel it as a way to solve their problem, as a way to approach their problems, because a lot of people don't understand what exactly is deep learning, how it works, how to initialize all of the variables, which are the hyperparameters per se. These all things are some limitations of deep learning as of now, and we hope by the time technology advances, people get to know more about what deep learning is, how artificial intelligence can be achieved through deep learning, they'll be more open to this, and all of these limitations will be laid off.

So, guys, uh, that's it from my side, and I hope you got to know a lot about deep learning interview questions which might help you in cracking the interviews and landing a great job as data scientists, machine learning engineers, or artificial intelligence engineers, as a matter of fact. And one important thing what I would like to say is that data scientist roles are somewhat, you know, industry-specific, or I would say, if you are working in healthcare, you should know about the healthcare industry too, rather than just knowing about the data and the numbers. So, if you're working in, suppose, imagery, so you should know what images, you should know what you're dealing with. So, a good knowledge of the particular industry which you're working for will also provide you a great advantage over other competitors. And since you know a lot of these stuff with this video, I'm sure you might be able to land a job, a great job in any of these industries.

And with this, we have come to an end to this full course on AI agents. If you enjoyed listening to this course, please be kind enough to like it, and you can comment on any of your doubts and queries. We will reply to them at the earliest. And do look up for more videos and playlists, and subscribe to Edureka's YouTube channel to learn more. Thank you for watching and happy learning.