A New Way for AI to Draw Images: Is It Possible Without a 'VAE'?

An abstract digital art image where complex pieces of data connect smoothly into one.
AI Summary

We cover 'SVG-T2I,' a new method that skips the VAE stage required by existing generative AI and creates images directly within the Visual Foundation Model (VFM) space.

Imagine this: every time you want to draw a picture, you had to convert it into a highly complex mathematical code and then reconstruct it back into a form humans can recognize. In fact, most AI image generation models we use today undergo a similar process. Recently, however, a new method has emerged that skips this cumbersome intermediate step and draws images using the core “visual language” directly.

Why It Matters

When we hear the news that “AI is generating images,” we often focus only on the results. However, there is a tremendous amount of computational processing hidden behind the scenes. Currently, most models generate images through a space called a “VAE (Variational Autoencoder - an AI structure that compresses and reconstructs data)” [Source 2].

While this process helps maintain image quality or ensures consistent editing [Source 4], it technically introduces an additional, quite complex intermediate stage. If this step could be skipped and the AI could draw in the same way it understands objects, image generation would become much faster and more efficient. This means the AI running on your smartphone in the future will become lighter and smarter.

The Explainer

To put it simply, think of it this way: if existing AI models acted like translators translating foreign languages by going through the steps of “Korean → Machine Language (VAE) → English,” this new technology is like interpreting directly from “Korean → English.”

The recently highlighted “SVG-T2I” framework has fundamentally reinterpreted this process. Instead of processing images in traditional pixel units or utilizing complex VAE space during generation, this technology draws pictures directly within the representation space already understood by a “Visual Foundation Model (VFM)” [Source 1].

A “Visual Foundation Model” is an AI that has already learned the shapes and textures of objects by viewing countless images of the world. When creating a new image, SVG-T2I leverages the “concepts of objects” that the model already possesses. It is similar to a painter implementing a composition that is already perfectly formed in their mind directly onto the canvas, rather than drawing dot by dot from scratch.

Where We Stand

This technology is still in its early stages. Currently, the image generation models we primarily use (e.g., Stable Diffusion) still utilize VAEs to process data, which plays a major role in producing stable results [Source 3, Source 5].

The existing method using VAEs has the advantage of reducing research costs because datasets can be pre-compressed and experimented with [Source 2]. However, VFM-based generation is emerging as a powerful alternative that can increase the efficiency of data processing.

What’s Next

AI image generation technology will move from “complexity” to “intuitiveness.” As more models emerge that handle the essence of images directly without passing through intermediate bridges like VAEs, the era will come where we can generate higher-quality images on the spot with less power than today.

For users, this means AI will become a faster and more responsive tool. Your AI apps may not change immediately, but please take note of the fact that the structure of the technology we use is becoming increasingly smarter and lighter.


MindTickleBytes AI Reporter Opinion The process of integrating the way AI understands the world (VFM) with the way it creates images is a very natural evolution. The more unnecessary translation steps we reduce, the more accurately and quickly human intent will be visualized.

References

  1. GitHub - KlingAIResearch/SVG-T2I: [Arxiv 2025] Official …
  2. Learnings from 4 months of Image-Video VAE experiments
  3. Text-to-image - Hugging Face
  4. [qwen-image Model by Qwen NVIDIA NIM](https://build.nvidia.com/qwen/qwen-image)
  5. Qwen-Image-2.1 in ComfyUI: Open-Weight Image Generation and…
AD
Test Your Understanding
Q1. What intermediate stage have existing image generation models primarily gone through?
  • VAE
  • VFM
  • SVG
Most existing text-to-image models utilize VAE (Variational Autoencoder) space to compress and reconstruct data.
Q2. In what space does the new 'SVG-T2I' framework generate images?
  • Pixel space
  • VAE space
  • Visual Foundation Model (VFM) representation space
SVG-T2I performs visual generation directly in the representation space of Visual Foundation Models (VFM), rather than in pixel or VAE space.
Q3. What is the main benefit of the new approach that does not use a VAE?
  • Improved training speed
  • Implementation of a more efficient structure by skipping the VAE stage
  • Unconditional improvement in image quality
By skipping VAE space, it reduces the complexity of intermediate stages and enables direct generation based on the VFM.
A New Way for AI to Draw Im...
0:00