Вход на сайт

Просмотр новости

Найдите то, что Вас интересует

Google is changing how it judges AI models for Android coding, updates list with Fable 5

Дата публикации: 08-07-2026 16:00:00

Google is changing how it tests AI models for Android coding following new rankings led by Fable 5.


Основное содержимое страницы с новостью.

Google is changing how it tests AI models for Android coding following new rankings led by Fable 5.

Google released the “Android Bench” earlier in the year as a glanceable ranking system for the best AI models. Those rankings are all based around the model’s ability to code for Android – a general ranking system it is not, but incredibly useful for developers.

Google says there are a couple of changes coming to how it ranks those AI models for Android development. The core benchmarking system is being changed to the standardized Harbor framework. As it stood, models like GPT-5.5, Claude Opus 4.7, and Gemini 3.1 Pro Preview were ranked based on a mini-swe-agent v1 benchmark tool developed for general use.

Switching tracks and opening the system up to the Harbor framework allows Android developers to use the same tools to analyze AI models for individual use cases. In fact, Google says it’s opening up the Android Bench to users willing to submit Android development tasks. Those will be used to evaluate how models handle each scenario. Developers are also being invited to share their benchmark evaluations.

From the beginning, we’ve valued an open and transparent approach, which is why we made our original methodology and test harness publicly available on
GitHub. You’ve asked for a way to provide feedback on our dataset, so now we’re taking collaboration a step further by giving you, the Android developer
community, a chance to shape Android Bench.

Android Bench adds Claude Fable 5

To top off the new approach, Google has refreshed the Android Bench list using the new framework for AI models. Each model between the closed-weight and open-weight variants has been reevaluated on the new testing bench.

Unsurprisingly, Claude Fable 5 sits at the top of the list with a rather comfortable lead. Google gave it a score of 84.5, a healthy 4 points ahead of GPT-5.5, ranked at 80.2. Claude Sonnet 5 is nearly 10 points lower than Fable 5. Anthropic’s power claims seem to be reasonable, and it’s worth noting that Fable 5 still has heavy restrictions put in place.

Most AI models that had a spot in the previous Android Bench rankings have received a new score, since the underlying rules have essentially changed. Here are the new rankings as Android Bench has them listed:

ModelScoreAvg LatencyAvg Cost
Claude Fable 584.58.0$133.2
GPT 5.580.215.7$138.3
Claude Sonnet 576.212.3$99.9
GPT 5.474.18.4$83.4
Gemini 3.1 Pro Preview73.710.6$87.4
Claude Opus 4.872.46.7$88.0
GLM 5.272.238.9$117.0
Gemini 3.5 Flash71.128.3$165.6
Kimi K2.7 Code70.431.8$48.1

Add 9to5Google as a preferred source on Google Add 9to5Google as a preferred source on Google

FTC: We use income earning auto affiliate links. More.

Схожие новости

#Наименование новостиТональностьИнформативностьДата публикации
1Gemini 3.5 Flash lands on Google’s Android coding rankings, but it’s 3x the cost for slower performance013.2212-06-2026
2Google изменила критерии отбора лучших ИИ для создания приложений под Android0510-07-2026
3Claude Opus 5 launches with similar performance as Fable 5 for ‘half the price’018.1224-07-2026
4You can thank AI for Google Keeps’s new large FAB on Android0304-06-2026
5Google will let websites opt-out of AI Mode & Overviews in Search0503-06-2026
6Google Finance now available as dedicated Android app0525-06-2026
7Google Is Cooking: Gemini 3.5 Pro Mogs Claude Fable 5 in Arena Leak0002-07-2026
8Android developer verification on track for September, ‘Verifier’ service will soon auto-install0518-06-2026
9Google makes AI Mode Pro visuals free ‘this summer,’ details Gemini tools for 2026 World Cup09.4509-06-2026
10Fable 5 just set a new AI freelance work performance record - but it can't replace humans yet3702-07-2026

Классификация: Пресс-релизы. Схожих патентов: 0. Схожих новостей: 10. Тональность: 0. Информативность: 11.25. Источник: 9to5google.com.