News

Meta’s new self-developed image model is here! There is no need to repeatedly adjust the prompt words. For the first time, Agent is introduced to automatically change the picture.

2 min read
Source: zhidx.com
Compiled by Zhidongzhi | Edited by Eggplant | Cheng Qian Zhidongzhi reported on July 9 that on July 7, Meta officially released the first multi-modal generative model developed by Meta Superintelligence Labs (MSL) - Muse Image. Meta also unveiled the video generation model Muse Video and demonstrated the videos it generated. However, the model is still in the testing phase and has not yet been officially launched. Muse Image is Meta's exploration of introducing Agent capabilities into the field of image generation. Unlike traditional AI drawing tools that directly generate pictures based on prompt words, Muse Image is more like an AI agent that can not only generate and edit pictures, but also independently search for information, call code tools according to task needs, and continuously optimize the generated results. Meta has released many pictures and videos generated by Muse Image and Muse Video, including pictures of moving people, animals and people in the same frame, as well as complex scenes where people are combined with food, scenery and other elements. The images below are all images generated by Muse Image, covering scenes of characters wearing sunglasses eating burgers, characters and animals sharing a table, and scenes of characters blending with stone pillars, blue sky and white clouds and other scenery. ▲People, animals and landscape images generated by Muse Image (Source: Meta) The picture below is a video clip of a person pouring water in the kitchen generated by Muse Video. ▲Character video generated by Muse Video (Source: Meta) According to foreign media TechCrunch, Muse Image was previously codenamed "Mango" internally and belongs to the Muse series of AI models being built by Meta. Alexandr Wang, head of the Meta Super Intelligence Laboratory, said that Muse Image has Agentic capabilities and can work together with the Muse Spark large language model to complete reasoning, search and planning capabilities before generating images. ▲Alexa