Transcription
As the field of artificial intelligence rapidly develops, everyone, from small startups to large companies, is trying to occupy their niches in this area and deliver a quality and useful product. The community cannot keep up with testing and bookmarking new products, and updates come almost daily in all directions. In this issue, we have gathered unusual and worthy solutions that various companies offer on the market. Make yourself comfortable, we are starting. Do you know what most people dislike who want to test something, something new, cool - it's paid subscriptions. But there are services that help do this for free, thereby solving this problem. Seriously, I think someone definitely doesn't want you to know about Tulify AI. This is not just a catalog, it is your personal free pass to the world of premium neural networks. Forget about endless registrations and trial versions for a couple of requests or a few days. Here is real access right now. Here's how it works. Lifai has gathered the best tools and neural networks for design, coding, text, video, and music in one place. And the most important thing is that it shows all their limitations on the free tariff immediately without surprises after an hour of work. Personally, I saved a lot of money this month alone by testing 10 different neural networks through Tifay. Want to compare new video generators? No problem. Play with an advanced agent or assistant, please, everything is real, without a catch and without spending a cent. Simplest interface, localization into 10 languages, convenient filtering on the site. There is even an option to customize the "My Tools" window. If you like to test new things first or you need to compare several options and you value your time, then register using the link in the description absolutely for free. Start testing any tools right now. Lifai is not just convenient. It is a game-changer for anyone who wants to stay up-to-date with trends without unnecessary expenses.
After a successful start to the summer with the release of powerful free open-source language and coding-focused AI models that not only matched but in some cases surpassed their closed-source US competitors, a team of researchers called QuH Team is back in action. They have released a new highly-rated open-source image generator model. The link was left in our Telegram channel in the pinned messages. Quen Image stands out from many other generative image models due to its emphasis on accurately rendering text within visual elements, an area where many competitors still struggle. The model supports both alphabetic and logographic writing and performs particularly well with complex typography, multi-line layouts, paragraph-level semantics, and bilingual content, such as English and Chinese. In practice, this allows users to create content like movie posters, presentation slides, shop displays, handwritten poems, and stylized infographics with clear text that matches their prompts.
You know, the data output measures in Quen Image include a wide range of real-world use cases: marketing and branding, bilingual posters with brand logos, stylized calligraphy, and unified design motifs. Presentation design, slides with consideration for heading hierarchy layout and corresponding thematic visuals. Education, creating learning materials with diagrams and accurate translation of educational text. Retail and e-commerce, storefronts where product labels, signage, and surrounding environments need to be clearly visible. Creative content, handwritten poems, scene descriptions, anime-style illustrations with embedded text. Users can interact with the model on the Quenchata website by selecting the image generation mode using buttons below the prompt input field. However, initial tests have shown that text-to-prompt alignment is not much better than Medpney, a popular proprietary AI-powered image generator from the eponymous American company. To my great disappointment, even after several attempts and rephrasing the prompt in the Quen chat, I encountered numerous errors in prompt understanding and text rendering accuracy. However, Mid Journey only offers a limited number of free generations, and more requires a subscription, unlike Quen Image, which, thanks to its open-source license and weights published on Hugging Face, can be used for free by any enterprise or third-party provider. Quen image is distributed under the Apache 2.0 license, which allows its use for both commercial and non-commercial purposes, as well as its distribution and modification. At the same time, derivative works require attribution and inclusion of the license text. This could make it attractive to companies looking for an open-source image creation tool for internal or external materials, such as flyers, advertisements, announcements, newsletters, and other digital communications. But the fact that the model's training data remains strictly confidential, as with most other leading AI image generators, may deter some companies from using it. QuH, unlike Adobe Firefly or OpenAI's own GPT-4 image generation system, for example, does not offer compensation for commercial use of its product. That is, if a user is sued for copyright infringement, Adobe and OpenAI will assist them in court.
The efficiency of Quen Image is backed by an extensive training process based on progressive learning, multimodal task alignment, and meticulous data processing. According to a technical paper published by the research group, the training corpus includes billions of image-text pairs taken from four sources: natural images, human portraits, artistic and design content, such as posters and UI mockups, and synthetic text data. The Quen team did not specify the size of the training data corpus beyond billions of image-text pairs. They provided an approximate breakdown by percentage of each content category: Nature 55%, Design UI, Posters, Illustrations 27%, People, Portraits, Human Activities 13%, Synthetic Text Rendering Data 5%. Notably, Quen emphasizes that all synthetic data was generated in-house and no images created by other AI models were used. Despite a detailed description of processing and filtering steps, the documentation does not specify whether any data was licensed or sourced from publicly available or private datasets. Unlike many generative models that exclude synthetic text due to the risk of noise, Image uses strictly controlled synthetic rendering pipelines to improve character coverage, especially low-frequency Chinese characters. It uses a curriculum-style strategy. The model starts with simple images with captions and non-textual content. Then it moves to text-heavy scenarios with layout considerations, multi-language rendering, and dense paragraphs. This gradual mapping is shown to help the model generalize scenarios and formatting types. Quen combines three key modules: 2.5 VL, which extracts contextual meaning and guides generation with system prompts; the QuH encoder, trained on high-resolution documents and real-world layouts, handles detailed visual representations, especially fine or dense text; and MMD IT, which is the foundation of the diffusion model, enabling joint learning on images and texts. Together, these components allow Quen Image to effectively handle image analysis, generation, and precise editing tasks. The QuH team emphasizes openness and community collaboration in releasing the model. Developers are encouraged to test and refine it, submit pull requests, and participate in the leaderboard evaluation. Feedback on text rendering, editing accuracy, and multilingual use cases will help improve future versions. The stated goal is to reduce the technical barriers to visual content creation. The team hopes that Mage will become not just a model, but a foundation for further research and practical application across various industries.
Google has just released one of its most powerful AI tools. And it's not a demo. It works, it's fast, and it's built to handle tasks that were previously thought to be only manageable by teams of people. Juls can dive into large-scale projects, understand every detail, and solve problems before you even notice them, plan your next steps, and continue working in the background while you focus on something else. It uses the same cutting-edge AI that powers Google's most advanced systems. And the way it integrates into real-world workflows will change how people create, develop, and bring ideas to life. Juls's idea from the beginning was simple: to provide developers with an AI agent that can work in the background, handle routine work, and allow you to focus on more important tasks. And now with version 2.0, they've added many new tools. Pull request integration, AI code review, environment snapshots, even built-in web surfing for documentation lookup. All on one platform without needing to switch between five different applications across a dozen browser tabs. We left the link in our Telegram channel in the pinned messages. Now, when Google talks about an AI agent, they mean it. When you give Juls a task, it doesn't just return a code snippet; it clones your entire codebase into a secure Google Cloud virtual machine, processes the full context of your project, and can continue performing tasks in the background while you focus on other work. It's designed to be both reactive and proactive. You can give it a task and let it run, or configure it to automatically perform certain updates and checks. It will even explain the logic behind the changes it makes. And if you want, it can generate audio recordings so you or your team can quickly get verbal reports on the work done. One of the things they've refined here is pull request management. Usually, you're switching between your IDE, GitHub, and maybe some internal review tools. Every context switch takes time. Jules allows you to handle pull requests directly on the platform. You can open, review, and merge without leaving it. And if you're on a feature branch, merging with the main branch takes just a few clicks. At first glance, this might not seem like a big deal, but for teams that merge branches frequently, it saves a lot of time. Another big update is environment snapshots. This will save a ton of time if you've ever had to stop a project midway and then spend half a day reconfiguring everything. Juls allows you to save the exact state of your development environment, VM configuration, dependencies, settings, so you can freeze them, test a new library, switch to another project, and then come back days or weeks later and restore everything exactly as it was. The new interactive task planning allows for real-time adjustments. In a typical agile workflow, priorities can change weekly, daily, and sometimes even hourly. Juls allows you to change task deadlines, dependencies, outcomes, and everything else directly on the platform. This is crucial for large teams that need to pivot quickly without disrupting their workflow. And there's also something that will seem too convenient once you get used to it: built-in web surfing for documentation lookup. If you're integrating an API or troubleshooting dependencies, Juls can pull up the latest official documentation, code examples, and guides without opening another browser window. You stay in your development environment, get what you need, and keep working. It's a small thing, but it helps you stay focused. Testing and previewing have also been improved. Dus works with Playwright for automated browser testing, so you can check functionality across different environments without manually launching multiple sessions. And before submitting, you can generate visual previews and screenshots to ensure everything looks right. This isn't just for frontend developers. If you're working on anything with a user interface, this will help catch visual issues before they go to production. Now, one of the most exciting features is what Google calls the Critic Agent. This is an AI-powered code review tool that uses reinforcement learning to identify errors, inefficiencies, and potential issues in real-time. It doesn't just point out syntax errors; it identifies performance bottlenecks, suggests optimizations, and flags logic that might break under certain conditions. If you've ever had a senior developer review your code line by line and tell you where you might go wrong, this is exactly that kind of feedback, but instant and without waiting for human review. And it's worth noting that Juls doesn't claim to replace an entire development team. Google has made that clear. It's more of a standalone assistant that can handle specific tasks autonomously. You can ask it to create and run unit tests, add a feature, update dependencies, or fix bugs, and it will get to work while you're doing something else. It's asynchronous, so you don't have to wait for a response like in a chat. GitHub integration is also very convenient. By connecting a repository, you can assign tasks directly from your existing development environment, and Juls will work with any branch you specify. All code changes it proposes will be presented for review. You can still make the final decision before merging, which is important for teams that need to maintain strict control over their codebase. Under the hood, it all runs on Gemini 25 Pro, Google's most advanced model to date. This is what allows Juls to tackle complex tasks, analyze large codebases, and perform multiple operations in parallel without getting derailed. But what about pricing? $2 gets you 15 daily tasks and three concurrent ones. $20 gets you 100 tasks and 15 concurrent ones. For $250, you get 300 daily tasks and 60 concurrent ones. And of course, priority access to new models as they become available. If you've been waiting for a coding assistant that truly becomes part of your workflow, not just an add-on gadget, then this is perhaps the closest we've seen to that from Google.
Researchers from Salesforce and the University of Southern California have developed a new method that allows computer agents to execute code while navigating graphical user interfaces, meaning they can write scripts while simultaneously moving a cursor or clicking buttons in an application, combining the best of both approaches to speed up workflows and reduce errors. This hybrid approach allows the agent to avoid brittle and inefficient mouse clicks for tasks that are better handled by code. The system, called CAT1, sets a new standard in key agent benchmarks, outperforming other methods while requiring significantly fewer steps to complete complex computer tasks. This advancement could pave the way for more robust and scalable agent automation, which has significant potential for real-world applications. Computer-using agents typically rely on Vision Language and Vision Language Action models to perceive screen images and perform actions, mimicking how a human uses a mouse and keyboard. While these GUI agents can perform various tasks, they often fail when executing long and complex workflows, especially in applications with many menus and options, such as office suites. For example, a task that involves finding a specific table in a spreadsheet, filtering it, and saving it as a new file might require a long and precise sequence of GUI manipulations. This is where brittleness comes in. In such scenarios, existing agents often face the challenge of visual perception ambiguity, such as needing to distinguish between visually similar icons or menu items, and the likelihood of making some error in the long run," the researchers write in their paper. One wrong click or a misunderstood UI element can derail the entire task. To address these issues, many researchers have focused on augmenting GUI agents with high-level planning. These systems use powerful reasoning models, such as OpenAI's GPT-4, to break down a user's global goal into a sequence of smaller, actionable subtasks. While this structured approach improves performance, it doesn't solve the problem of navigating menus and clicking buttons, even for operations that can be performed more directly and reliably with a few lines of code. To overcome these limitations, researchers created CA1, Computer Using Agent with Coding as Actions, a system designed to combine the intuitive, human-like capabilities of a graphical user interface with the precision, reliability, and efficiency of direct system interaction through code. The system is structured as a team of three specialized agents that work collaboratively: an orchestrator, a programmer, and a GUI operator. The orchestrator acts as the central planner or project manager. It analyzes the user's overall goal, breaks it down into subtasks, and assigns each one to the most suitable agent for the job. It can delegate internal operations like file management or data processing to the programmer, who writes and executes scripts in Python or Bash. For frontend tasks requiring button clicks or navigation through visual interfaces, the GUI operator, a VLM-based agent, is used. This dynamic delegation allows CoaK1 to strategically bypass inefficient GUI sequences in favor of reliable, one-shot code execution where appropriate, while still leveraging visual interaction for tasks where it is necessary, the paper states. The workflow is iterative. After the programmer or GUI operator completes a subtask, they send a summary and a screenshot of the current system state to the orchestrator, who then decides on the next step or concludes the task. The programmer agent uses an LLM to generate its code and sends commands to a code interpreter to test and refine its code over several rounds. Similarly, the GUI operator uses an action interpreter that executes its commands, such as mouse clicks, text input, and returns the resulting screenshot, allowing the operator to see the outcome of its actions. The orchestrator makes the final decision on whether to proceed with the task or stop it. Researchers tested CA1 on OS World, a comprehensive benchmark, and observed the greatest performance gains in categories where code execution offers a clear advantage, such as in operating system-level tasks and multi-step workflows. For instance, consider an OS-level task: finding all image files in a complex folder structure, resizing them, and then compressing the entire directory into a single archive. A purely GUI-based agent would have to perform a long and unreliable sequence of clicks and drags, opening folders, selecting files, and navigating menus, with a high probability of error at each step. CoaK1, on the other hand, can delegate this entire workflow to its programmer agent, which can execute the task with a single, robust script. In terms of concurrency, the entire process takes it approximately 10 steps, compared to OpenAI's same QA which requires 40 steps. Researchers identified a clear pattern: tasks requiring more actions are more likely to remain uncompleted. Reducing the number of steps not only speeds up task completion but, more importantly, minimizes the probability of error. Thus, finding ways to combine multiple GUI steps into a single code task can make the process more efficient and less error-prone. Despite high test results, enterprise environments are far more complex. They involve legacy software and unpredictable user interfaces. This raises important questions about reliability, security, and the need for human oversight. The primary challenge lies in the orchestrator agent making the correct choice when dealing with an unfamiliar application. According to the developers, to make agents like CAK-1 robust to non-standard enterprise software, they need to be trained using feedback in realistic, simulated environments. Ultimately, in the foreseeable future, ambiguity resolution will likely require human involvement. Responding to a question about how to handle vague user requests, which is also mentioned in the paper, a developer suggested a phased approach. "I think you need to involve a human to start," he noted. While some tasks may become fully automated over time, for high-stakes operations, human verification will still be crucial. Some critical tasks will always require human approval.
Thank you for watching until the end. Subscribe to YouTube and Telegram channels, and also write your opinion in the comments. Thank you for your attention and see you in new releases.