An AI Model Tries to Recreate Super Mario Bros through Generated Videos | When did Roblox become popular | PC gaming vs console gaming statistics | Gaming industry worth worldwide | Turtles AI

An AI Model Tries to Recreate Super Mario Bros through Generated Videos
MarioVGG explores the possibilities of automatic gameplay generation, showing the first steps towards creating games through AI
Editorial Team6 September 2024

 

In our recent deep dive into the potential of AI-based video game generation, we explored Google’s GameNGen AI model that managed to create a playable version of Doom. Now, a new model called MarioVGG, developed by researchers at AI company Virtuals Protocol, attempts to replicate Super Mario Bros. gameplay through AI-generated videos. The results, although still far from perfect, open new avenues for the future of video games.

Key Points:

  • AI for the Game Generation: MarioVGG attempts to create Super Mario Bros. gameplay videos based on input data and in-game images, mimicking the physics and dynamics of the game.
  • Techniques and limitations: The model, which uses convolution and denoising, still has many glitches and a processing speed that is too slow for real-time gameplay.
  • Model Training: MarioVGG was trained on over 737,000 game frames, with a limited focus on two main inputs: “run right” and “run and jump”.
  • Future prospects: Despite current limitations, the model represents a step towards creating reliable and controllable video game generators, with possible future applications in game development.

MarioVGG is an AI model designed to generate Super Mario Bros. gameplay videos, based on input data and pre-existing video sequences. Unlike other previous attempts, such as the aforementioned AI GameNGen from Google which we wrote about in depth in a previous article of ours, MarioVGG focuses on the simulation of a specific title, trying to replicate its physics and game dynamics. To train the model, the researchers used a public dataset containing 280 Super Mario Bros. levels, with over 737,000 individual game frames. These frames were preprocessed in blocks of 35 to allow the model to learn the effects of the inputs, focusing on two main actions: running right and jumping while running. Despite this simplification, the system encountered significant difficulties, especially in handling complex inputs such as jumps with mid-air adjustments, which were discarded to avoid noise in the training data.

After approximately 48 hours of training on a single RTX 4090 graphics card, the researchers used a convolution and denoising process to generate new video frames starting from a static game image and textual input. The generated sequences are short, but the last frame of each sequence can be used as the starting point for a new one, thus allowing the creation of gameplay videos of potentially unlimited length. However, visual quality remains limited, with output frames scaled to a resolution of 64x48 pixels, much lower than the original resolution of the NES. Furthermore, the model is still far from generating real-time video: a single RTX 4090 takes six seconds to create a sequence of six frames, representing just over half a second of video, making the model’s use for interactive video games impractical. Despite these limitations, MarioVGG has proven to be able to learn game physics from training data, simulating behaviors such as Mario falling when running off a cliff or stopping movement in the presence of obstacles. However, the model is not free of errors: in some cases, MarioVGG ignored user prompts, generating visual glitches or absurd behavior, such as Mario transforming into a Cheep-Cheep during a jump. The researchers suggest that longer training on more diverse data could improve the model’s ability to simulate gameplay more accurately.

Despite its limitations, MarioVGG represents an interesting step towards creating AI models capable of autonomously generating video games, a field that could revolutionize game development in the future.