New horizons in the generation of images ai | Ai generator | | Dall-e pretrained model | Turtles AI

New horizons in the generation of images ai
Zhipu Technology presents the Cogview3 and Cogview3-Plus models, now open source and ready to transform the panorama of digital creativity
Editorial Team16 October 2024

 

Innovation in the generation of images: Zhipu Technology’s Cogview3-Plus model marks a significant step forward for Text-to-IMAGE technology.

Key points:

  • Cogview3 and its advanced Cogview3-Plus version are now Open Source.
  • The generation process includes three phases, from lower to high resolution.
  • Cogview3 significantly improves performance compared to existing models.
  • The new models open to future applications in the field of digital creativity.

Zupi Technology has recently made available to the public his latest innovation in the field of the generation of images assisted by AI, represented by the Cogview3 and Cogview3-Plus models. These tools, available through the "Zhipu Qingyan" app, mark an important evolution in the panorama of Text-to-IMAGE technology, allowing users to explore new ways of artistic creation. Cogview3 uses a waterfall diffusion approach, which is divided into three phases: initially generates a low resolution image of 512x512 pixels, then the image is refined through a diffusion process that leads to a resolution of 1024x1024 pixels and finally, finally, to a further iteration produces a high definition image of 2048x2048 pixels. This methodology recalls the work of an artist who gradually refines his work, improving the final visual quality. The tests highlighted how Cogview3 exceeds the performance of the current Open Source Standard in the sector, SDXL, achieving considerably higher results of 77%. In addition, the rapidity of inference of the new model is ten times faster than SDXL, testifying to the optimization work carried out by Team Zupu. The next version, Cogview3-Plus, brings significant innovations with it, such as the integration of the dit framework and the adoption of the zero-snr diffusion planning planning, which further improve overall performance. In addition, the implementation of a joint attention mechanism for text and image allows to optimize costs and resources, creating a balance between efficacy and efficiency. The new model uses a 16 -dimensional latent space, opening promising roads for future developments in the generation of images. For developers and researchers interested in experimenting with these technologies, Zhipu Technology has made accessible the repositories of the source code, thus facilitating the progress in the sector. The introduction of the Cogview3 models expands the application potential of Text-to-IMAGE technology, with implications ranging from personal artistic creation to commercial, educational and entertainment sectors. In this context, the generation assisted by artificial intelligence could become increasingly common, allowing an increasing number of people to express their artistic ideas.

These developments lay the foundations for a future in which human creativity and technological innovation collaborate in increasingly synergistic ways.