Search
Benchmarking AI agents on real design tasks. V0.1 includes 236 Framer canvas challenges measuring model quality and efficiency across layout, design understanding, edits, and web design workflows. Unlike single-metric benchmarks, it reflects the work required to create and maintain high-quality websites at scale, offering a more holistic comparison of model strengths.
V0.1
July 9, 2026
The primary measure of a model is our most challenging benchmark. Given a user-generated, non-functional site navigation design that only defines alignment, spacing, and colors. The task is to turn it into a responsive, functional nav with overlaying mega menus on desktop and a clipped sliding drawer on mobile.
Fable 5
100%
5m
22
$4.99
Strong result, but costly
Opus 5
100%
3.5m
25
$2.63
Fable 5 result, faster and cheaper
Opus 4.8
63%
5m
20
$2.98
Misses critical functionality
Sonnet 5
63%
6m
27
$1.56
More steps, slower, costlier, got stuck
GPT 5.6 Sol
100%
2m
15
$1.37
Strong result
GPT 5.6 Terra
100%
1m
14
$0.658
Strong result, cheapest and fastest
GPT 5.5
90%
2m
18
$1.32
Lower quality at similar cost
The results are mixed rather than one-sided. Sol is strongest on task completion, Terra is efficient, and Fable leads several consistency-related categories. Each model shows different strengths across the benchmark.
Responsive Design
75.7%
In 25.1M · Out 52k
70.3%
In 22.1M · Out 47k
72.5%
In 29.1M · Out 63k
74.1%
In 29.2M · Out 78k
68.6%
In 39.1M · Out 122k
73.7%
In 24.0M · Out 89k
71.7%
In 24.8M · Out 86k
Interactions
95.8%
In 9.4M · Out 18k
75.3%
In 7.2M · Out 14k
70.6%
In 9.2M · Out 24k
71.5%
In 11.2M · Out 39k
90.6%
In 19.7M · Out 47k
86.5%
In 12.7M · Out 70k
81.9%
In 16.0M · Out 60k
Layout & Spacing
75.5%
In 17.9M · Out 65k
79.6%
In 13.1M · Out 47k
76.9%
In 14.2M · Out 67k
78.4%
In 18.7M · Out 66k
74.3%
In 25.4M · Out 85k
74.1%
In 15.5M · Out 87k
80.0%
In 18.6M · Out 87k
Components & Templates
79.3%
In 18.2M · Out 33k
80.3%
In 10.2M · Out 19k
85.3%
In 16.2M · Out 33k
92.5%
In 21.2M · Out 58k
85.0%
In 30.2M · Out 70k
80.7%
In 14.1M · Out 78k
84.6%
In 18.2M · Out 73k
CMS
86.8%
In 19.7M · Out 45k
74.1%
In 13.7M · Out 29k
89.0%
In 22.2M · Out 62k
91.1%
In 27.9M · Out 80k
83.8%
In 28.6M · Out 72k
82.5%
In 21.6M · Out 101k
84.7%
In 23.0M · Out 96k
Reuse & Consistency
79.2%
In 8.6M · Out 14k
66.1%
In 3.8M · Out 9.2k
82.7%
In 7.8M · Out 21k
87.5%
In 5.9M · Out 19k
66.0%
In 5.6M · Out 42k
66.0%
In 5.6M · Out 42k
86.6%
In 8.2M · Out 43k
Shaders
100.0%
In 1.1M · Out 1.9k
90.0%
In 1.2M · Out 1.7k
75.7%
In 1.2M · Out 2.3k
100.0%
In 1.2M · Out 1.8k
100.0%
In 633k · Out 1.5k
100.0%
In 633k · Out 1.5k
100.0%
In 878k · Out 1.8k
A good score requires the model to
Create a component, create the right number of navigation visual variants, and assign them to the correct breakpoints on the main webpage.
Avoid syntax errors, use tools efficiently
Add the required best-practice interactivity: add relative overlays, create the visual states for the open and closed drawer.
A great score requires the model to
Rhythm: Understands and preserves the spacing rhythm of the original design
Adherence: Maintains exact visuals & layout while executing complex component reparenting and reorganisation
Micro-interactions: Shows UI / UX tastefulness via hover animations (e.g. adds an animation to the overlay trigger chevrons so that they rotate cleanly when the overlay is visible)
In this benchmark task, a perfect score (100%) means the result can be published without any follow-up fixes.
Create personal portfolio
Build startup site
Launch landing page
Start company blog

