Dust is the first technique to train Transformer AI using 'zeroth-order optimization' instead of traditional backpropagation, showing potential to revolutionize computational efficiency.
Have you ever wondered how the AI chatbots we use every day learn to speak so intelligently? Until now, the most common way AI learns has been ‘backpropagation’ (a method of learning by passing errors backward). It’s similar to a student taking an exam and tracing their mistakes from the last question back to the first to identify and correct them. However, this method has long suffered from chronic issues: as AI models grow larger, they require enormous computational power, and the learning process itself is highly complex.
Recently, however, a research result suggesting that AI can be trained without using backpropagation at all has garnered attention. This new learning method is called ‘Dust.’
Why is this important?
As AI technology advances, we demand larger and more complex models. But as models grow, backpropagation hits a massive wall known as ‘computational cost.’ While backpropagation has long been the standard for deep learning, its limitations—such as high computational requirements, weight transport problems, and issues with stalled or erratic learning—have been widely documented.
What if AI could learn on its own without the complex bridge of backpropagation? The time and electricity required to train AI would decrease, allowing more efficient artificial intelligence to be brought to the world faster. Beyond just being a new technique, Dust may be the key to solving the ‘bottleneck’ of AI training.
How does it work?
If backpropagation is a ‘precise tutoring session where you read the textbook backward to correct your mistakes,’ what is Dust like?
In simple terms, it’s like an ‘intuitive experiment.’ Imagine having to assemble a complex machine: instead of following the instructions step-by-step, you randomly swap out parts and simply check the ‘actual results’ to see if the machine works better. In technical terms, this is called ‘zeroth-order optimization.’ When pretraining Transformer AI, instead of using traditional backpropagation, Dust solves the problem using forward evaluation and stochastic gradient descent (SGD, a method that gradually reduces errors based on data).
Rather than flipping the entire process on its head, it chooses to observe the results and directly make incremental corrections. With sufficient computational scale, Dust demonstrates performance similar to, or sometimes surpassing, traditional backpropagation methods.
Current Status
Dust is already proving its potential beyond simple ideas through real-world experiments. It is particularly significant as the first zeroth-order optimization method used to pretrain a Transformer model.
What’s even more surprising is its efficiency. Research results indicate that Dust is approximately 1,000 to 10,000 times more computationally efficient than EGGROLL, an existing non-backpropagation training method. Of course, it is still in the early stages of development and is currently undergoing the process of proving its performance through massive-scale computations.
What comes next?
The emergence of Dust has the potential to change how we build AI. Studies like Dust suggest new paths to bypass the inherent limitations of backpropagation, such as vanishing gradients (where learning signals disappear as they propagate backward) or exploding gradients.
If AI becomes more efficient in its learning, high-performance AI model training—once the exclusive domain of giant corporations—could become more accessible to individual researchers. Whether Dust can completely replace the sophistication backpropagation has built up over decades remains a task that requires more data and large-scale experimentation. What is certain, however, is that the world of AI training is moving past the era of relying solely on backpropagation, evolving into more diverse and efficient methods.
AI Opinion
Dust is a bold attempt that could shake up the existing training paradigm. If efforts to maximize efficiency by breaking out of the massive framework of backpropagation succeed, the pace of artificial intelligence development will likely be far faster than we imagine. While there is still a long way to go, it is clear that the way AI teaches itself is becoming lighter and smarter.
References
- A claimed way to pretrain transformers without backpropagation
- dust/README.md at main · qlabs-eng/dust · GitHub
- Navigating beyond backpropagation: on alternative training … - Springer
- Pretraining with Random Noise for Fast and Robust Learning - arXiv:2405.16731
- Samip on X: “Backprop has been the only credit assignment …”
- Reinforcement Learning
- Zeroth-order optimization
- Transfer Learning
- High computational resource requirements
- Vanishing and exploding gradient problems
- Excessively fast learning speed
- 100–500 times
- 1,000–10,000 times
- 2 times