The 'Agents on Rails' benchmark, which measures how well AI can implement complex features in actual Ruby on Rails projects, recorded a 35% success rate for the top model, proving its potential for real-world application.
Imagine this: You wake up in the morning and tell your AI assistant, “Add a user registration feature to our website today and check for any related security issues.” While you enjoy a cup of coffee, the AI writes the complex code, runs tests on its own, and reports back, “All tasks completed.”
| This was something we only saw in science fiction movies just a few years ago, but now we are one step closer to this future. Just how well are current AI models performing actual tasks on behalf of developers? We’ll take a look at the reality through the results of the ‘Agents on Rails’ benchmark, recently released by the Ruby on Rails Foundation and Evil Martians. [Source: Agents on Rails: the first benchmark report, [Source: Rails Foundation launches an AI coding agent benchmark for Ruby on Rails | daily.dev](https://daily.dev/posts/rails-foundation-launches-an-ai-coding-agent-benchmark-for-ruby-on-rails-shazaa4gk)] |
Why Is This Important?
While many AI models have promoted their coding abilities, real-world corporate projects are much more complex and demanding. Most existing benchmarks have only tested very short and simple snippets of code.
‘Agents on Rails’ is important because it is a ‘real-world test.’ It takes the actual code of a project called ‘Writebook’ that developers use and tasks the AI with solving practical challenges like bug fixes, security audits, and new feature additions. [Source: Agents on Rails: the first benchmark report, Source: Agents on Rails benchmark: model picks by cost and score] In other words, these results serve as a ‘practical report card’ that tells us how much we can trust and rely on AI if we were to introduce it into our work environment today.
Easy to Understand
It’s easy to understand this benchmark with an analogy.
Simply put, if previous AI performance measurement methods were like taking an ‘elementary school-level English vocabulary test,’ ‘Agents on Rails’ is like a ‘practical ability assessment’ where you have to join a company in an English-speaking country, write reports, and collaborate like a new hire.
The AI agent is like a new employee who has just joined the company. In the Stage 1 test, it was assigned very short and independent tasks (finding bugs, solving security issues, etc.), and in the Stage 2 test, it was tasked with performing the ‘entire feature implementation process’ just like an actual developer would. [Source: Agents on Rails: Stage 2. Can a model ship a feature?, Source: Agents on Rails: We ran 8 models against 21 atomic tasks to …]
In the recently announced Stage 2 results, the ‘GPT-6 Astra’ model showed the best performance with a success rate of 35%. You might feel, “Huh? That’s lower than I thought.” However, the fact that an AI can successfully complete 35% of complex, real-world tasks on its own means that it is at a level where it can dramatically increase work efficiency if an experienced developer reviews and corrects its work alongside it.
Current Situation
Currently, ‘Agents on Rails’ is rigorously testing 8 major AI models. [Source: Rails Releases First AI Coding Agents Benchmark]
-
Performance of Top-tier Models: In the Stage 1 test, ‘Claude Opus 5’ recorded a surprising success rate of 92%. [[Source: Agents on Rails: the first benchmark report Vuink.com](https://vuink.com/post/eholbaenvyf-d-dbet/2026/8/13/agents-on-rails-the-first-benchmark-report)] - Various Options: ‘Kimi K3’ proved its efficiency by delivering 90% of the performance of top-tier models at half the cost, and ‘GPT-5.6 Luna’ drew attention for being the most cost-effective. [Source: Rails Releases First AI Coding Agents Benchmark]
- Limitations: However, as seen in the Stage 2 test where entire features must be implemented, AI still needs improvement to fully understand the overall context of real-world projects and complete code without errors.
What’s Next?
| AI coding agents will become even smarter in the future. The Rails Foundation plans to continue evolving the benchmark by comprehensively evaluating not just the success rates of models, but also how well they reflect the latest development patterns and whether token costs (the unit cost incurred when AI processes data) are appropriate. [[Source: Rails Foundation launches an AI coding agent benchmark for Ruby on Rails | daily.dev](https://daily.dev/posts/rails-foundation-launches-an-ai-coding-agent-benchmark-for-ruby-on-rails-shazaa4gk)] |
What readers should pay attention to is not the simple score, but the ‘trend.’ We are moving from the stage where AI simply knows grammar to the stage where it directly implements features that create actual business value. The moment the 35% success rate rises to 50% or 70% in the near future, the way we work will change completely.
MindTickleBytes AI Reporter’s View
This benchmark proves that AI is growing beyond a ‘coding assistant’ into a ‘colleague.’ While the 35% figure is not perfect, it is more hopeful than any other result in that AI has begun to understand and execute the actual workflows of developers.
References
- Agents on Rails: the first benchmark report
- Agents on Rails: The LLM Benchmark Project
- Agents on Rails: Stage 2. Can a model ship a feature?
- Agents on Rails: Grok 4.6, GLM 5.3, Gemini 3.7 Flash, and Opus 4.8
-
[Agents on Rails: the first benchmark report Vuink.com](https://vuink.com/post/eholbaenvyf-d-dbet/2026/8/13/agents-on-rails-the-first-benchmark-report) -
[Rails Foundation launches an AI coding agent benchmark for Ruby on Rails daily.dev](https://daily.dev/posts/rails-foundation-launches-an-ai-coding-agent-benchmark-for-ruby-on-rails-shazaa4gk) - Agents on Rails benchmark: model picks by cost and score
- Agents on Rails: We ran 8 models against 21 atomic tasks to …
- What the First Rails Agent Benchmark Tells You, and What It …
- Rails team’s first “Agents on Rails” benchmark report: how well do models actually know Rails APIs?
- Rails Releases First AI Coding Agents Benchmark
- Writebook
- RailsApp
- CodeAgent
- 92%
- 35%
- 50%
- To measure AI's graphic processing capabilities
- To measure AI coding performance in an environment similar to actual work
- To test AI's writing skills