LiveCanvas AI web design benchmark: an exact prompt excerpt points to the real Claude Sonnet 4.6 page, showing its hero, weekly plan and benefits.
General

AI Page Builder for WordPress: Which Model, and at What Cost?

Lorenz

Lorenz

September 23, 2026

Join our Discord

Which AI model should you choose to build a WordPress page? How much will the generation cost, and how close will the result be to something you can use on your site? Those questions matter when you bring AI into a developer's workflow. A page needs to fit the design you already have, and you need to be able to change it afterwards.

We used LiveCanvas as an AI page builder for WordPress to explore those questions. Its AI Assistant let us send the same page brief to different models through OpenRouter, then work with the generated HTML inside the editor. On 23 September 2026, we tried 22 OpenRouter options on one staging site. Twenty-one produced a page; one endpoint was rate-limited on both attempts.

The fastest call, Qwen3 Coder, generated a six-section page in 22.334 seconds for $0.00821772, less than one US cent. That page matched the supplied theme and stayed within both viewport widths we checked. It also contained a product claim we would edit before publishing. That combination is the point of this experiment: a small generation cost can buy a useful first draft, with the developer still responsible for the finished page.

LiveCanvas AI benchmark: an exact prompt excerpt points to the actual Claude Sonnet 4.6 page, showing the hero and weekly meal plan.
One shared brief, 22 OpenRouter options. The prompt excerpt points to Claude Sonnet 4.6’s actual first result; each model’s recorded cost and generation time are compared below.

What does AI actually build inside LiveCanvas?

For this test, the AI wrote the frontend HTML for a page: its layout and copy, including a meal-plan preview and a working Bootstrap FAQ accordion. LiveCanvas placed that markup in the page we were editing. The WordPress theme continued to supply the site's header, footer and typography.

OpenRouter connects the Assistant to the selected model. Its unified API provides access to models from different companies through a common interface. Claude, GPT, Gemini and the other names in the tables are the models doing the generation. OpenRouter handles the request to the provider and records usage. This makes it possible to compare different options while keeping the same LiveCanvas workflow.

We began each attempt with a blank page and an empty section, chose the model in the Assistant and submitted the same brief using Go ask AI. Website Information and Page Text were switched off because all the site context for this controlled test was written into the prompt. Once generation finished, we saved the page and checked the frontend.

The generated page remains editable. You can change its text through LiveCanvas or open Edit HTML to inspect and adjust the markup. In the North result below, the text sits inside editable="rich" wrappers. The AI has supplied a starting layout that a developer can keep working on within the same page editor.

LiveCanvas HTML editor beside the generated Hearthline page
The generated HTML remains accessible inside LiveCanvas.

Can an AI-generated page match your existing WordPress site?

The site context made the brief useful. Our staging site was Hearthline, a demo family organisation app built with a Bootstrap 5 theme. It already used purple buttons on a cream background, Playfair Display headings and Inter body text. We put those decisions into the prompt, including the primary colour #7047a8, secondary colour #c38fca and body background #fbf8f2.

We also carried over the product facts. Hearthline described shared calendars, meal plans, routines and lists. The offer was 30 days free, then EUR 7 per household per month, with unlimited family members. Those details gave the models material for a meal-planning feature page that belonged to this particular site.

The page brief asked for a two-column hero followed by a five-day meal-plan preview and shopping list. Benefit cards and a short workflow explained the feature; four FAQs and a final call to action completed the page. The models had to use Bootstrap, preserve the existing header and footer, and make text editable in LiveCanvas. We supplied the path to one image already in the WordPress media library.

The screenshots show different compositions built from that shared direction. Familiar colours and typography carry through the results because the brief gives the model concrete design constraints and the theme provides the fonts. For your own site, the same idea applies: describe the page you need using the design and product facts you have already settled.

How much does an AI-generated WordPress page cost?

In this experiment, the paid page-generation calls ranged from $0.0022517 to $0.22352. GPT-6 Luna had the lowest recorded paid charge and GPT-6 Astra the highest. The successful free endpoints recorded $0.00 in OpenRouter; the Gemma request used a BYOK route, which is noted with the free results below.

Those figures pay for one model response. They describe the cost of generating the initial HTML, separate from the cost of a finished page. A production budget also has to cover the developer's review and any further requests, alongside the site's normal software and hosting costs. No new images were generated in these calls.

Time varied too. The fastest provider generation took 22.334 seconds; the slowest successful one took 363.144 seconds, just over six minutes. That gives a more useful expectation than promising that every model finishes in a few seconds. Choosing a model can change how long you wait as well as what the first draft costs.

Compare the first results

All 21 successful calls from 23 September 2026. Cyan bars show the recorded charge in USD. Magenta bars show provider generation time. Shorter bars mean a lower charge or a shorter wait.

Each model also has a layout check at 390 px and 1440 px. Eight pages fit both widths; 13 had 12 px of mobile overflow. This checks horizontal overflow only. Visual quality has no common score in this test.

Model and layout check
Recorded cost, USDCyan scale: $0 to $0.25
Generation timeMagenta scale: 0 to 400 seconds
Qwen3 Coder

Fits both tested widths
Cost (USD)$0.00821772
Generation time22.334 s
GPT-6 Sol

+12 px mobile overflow
Cost (USD)$0.038534
Generation time38.428 s
GPT-6 Luna

+12 px mobile overflow
Cost (USD)$0.0022517
Generation time47.557 s
Claude Opus 5.5

+12 px mobile overflow
Cost (USD)$0.150148
Generation time56.768 s
GPT-6 Astra

+12 px mobile overflow
Cost (USD)$0.22352
Generation time73.218 s
Gemini 3.8 Flash

+12 px mobile overflow
Cost (USD)$0.032172
Generation time82.938 s
MiMo V2.6 Pro

+12 px mobile overflow
Cost (USD)$0.00459012
Generation time91.861 s
Claude Sonnet 4.6

Fits both tested widths
Cost (USD)$0.116214
Generation time93.129 s
Gemma 4 31B IT

+12 px mobile overflow
Cost (USD)$0.00BYOK route
Generation time97.723 s
Gemini 3.1 Pro Preview

+12 px mobile overflow
Cost (USD)$0.18209
Generation time102.821 s
North Mini Code

Fits both tested widths
Cost (USD)$0.00Free endpoint
Generation time120.962 s
Dots3 Note Preview

Fits both tested widths
Cost (USD)$0.00Free endpoint
Generation time166.214 s
DeepSeek V4.1 Flash

+12 px mobile overflow
Cost (USD)$0.0041923
Generation time180.958 s
Qwen3.8 Omni Flash

+12 px mobile overflow
Cost (USD)$0.00608499
Generation time192.840 s
Kimi K2.5

+12 px mobile overflow
Cost (USD)$0.0135846
Generation time195.962 s
Grok 4.7

Fits both tested widths
Cost (USD)$0.0749456
Generation time209.388 s
Ling 3.0 Flash VL

Fits both tested widths
Cost (USD)$0.00Free endpoint
Generation time210.597 s
OpenRouter Free router

+12 px mobile overflow
Cost (USD)$0.00Routed to Nemotron
Generation time219.967 s
GLM 5.3 Flash

Fits both tested widths
Cost (USD)$0.01026135
Generation time240.867 s
Nemotron 3 Ultra

+12 px mobile overflow
Cost (USD)$0.00Free endpoint
Generation time298.243 s
MiniMax M2.5

Fits both tested widths
Cost (USD)$0.021603925
Generation time363.144 s

Both bar scales start at zero and stay fixed when sorted: $0 to $0.25 for cost, 0 to 400 seconds for time. Values are individual calls, before saving and review. A $0.00 charge has no filled cost bar. Gemma used a BYOK route, so its OpenRouter charge does not establish a zero total cost.

Laguna S 2.1 returned HTTP 429 twice. Its missing cost and generation time are excluded from the bars. OpenRouter Free router reached Nemotron and is a separate call, so the rows count tested options.


Which AI model would we use for this page?

For this particular brief, Qwen3 Coder would be our first option to try again. It was the fastest successful call, cost less than one cent and had no horizontal overflow at either tested width. Its claim that ingredients would appear automatically on the shopping list was unsupported by the brief, so the copy still needed attention.

If the smallest paid generation charge were the priority, GPT-6 Luna would be worth a look. It cost $0.0022517 and took 47.557 seconds of provider generation, although its page had a small mobile overflow. Among the free options, North Mini Code produced an editable page that fit both widths, with copy that needed revision.

These observations help narrow a choice for a similar page. A single brief cannot establish the best AI model for every WordPress project. We did not score visual quality on a common scale or repeat runs to estimate average speed. A developer working on complex changes may value different capabilities, and this experiment only measured initial page generation.

The results, model by model

The tables group the shortlist into frontier models, default candidates and open weight options, followed by the lower-cost and free categories. Each row shows OpenRouter's recorded provider generation time and actual usage charge. These are individual calls, excluding briefing, saving and review.

Output tokens include reasoning tokens when the provider reports them that way. The reasoning column is a subset of output, so it should not be added again. We kept the default reasoning settings and allowed provider routing. That reflects this Assistant session, with the exact provider and model identifiers preserved in the data below.

The 21 successful calls totalled $0.888410305 in recorded usage. Laguna returned HTTP 429 on its initial attempt and one retry. We found no matching generation metrics for those failed attempts in the activity view, so their token counts, costs and generation time remain unavailable.

Frontier models

The frontier calls generated pages with the same supplied design context. Their costs ranged from $0.0749456 to $0.22352 in this run. This page brief does not test large refactors or end-to-end software development.

Model Input tokens Output tokens Reasoning within output Generation, s Actual cost, USD
Claude Opus 5.5 1,322 7,243 433 56.768 $0.150148
GPT-6 Astra 822 4,306 197 73.218 $0.22352
Gemini 3.1 Pro Preview 871 15,029 10,228 102.821 $0.18209
Grok 4.7 2,084 15,207 11,747 209.388 $0.0749456

Default candidates

Sonnet and Sol both generated the requested page. Sol used fewer output tokens and completed generation sooner in this run. An iteration test would be needed before calling either the best everyday model for every LiveCanvas task.

Model Input tokens Output tokens Reasoning within output Generation, s Actual cost, USD
Claude Sonnet 4.6 913 7,565 0 93.129 $0.116214
GPT-6 Sol 822 3,689 825 38.428 $0.038534

Open weight models

Qwen3 Coder was the fastest successful call in this set. MiMo and DeepSeek also produced pages for less than one US cent each. MiniMax and GLM spent much more of their output budget on reasoning.

Model Input tokens Output tokens Reasoning within output Generation, s Actual cost, USD
MiMo V2.6 Pro 844 4,854 546 91.861 $0.00459012
Kimi K2.5 828 5,872 1,003 195.962 $0.0135846
Qwen3 Coder 846 4,462 0 22.334 $0.00821772
DeepSeek V4.1 Flash 860 9,695 4,660 180.958 $0.0041923
MiniMax M2.5 835 17,798 13,970 363.144 $0.021603925
GLM 5.3 Flash 829 20,274 15,759 240.867 $0.01026135

Maximum savings candidates

GPT-6 Luna had the lowest recorded cost among the paid calls. Gemini 3.8 Flash cost $0.032172 here. The category describes the shortlist we tested; it is not a ranking of the final bill.

Model Input tokens Output tokens Reasoning within output Generation, s Actual cost, USD
GPT-6 Luna 822 4,339 873 47.557 $0.0022517
Gemini 3.8 Flash 871 8,405 2,362 82.938 $0.032172

Budget multimodal

Qwen3.8 Omni Flash generated HTML from the same text prompt. We did not send an image, audio or video input in this test, so this result does not measure its multimodal capabilities.

Model Input tokens Output tokens Reasoning within output Generation, s Actual cost, USD
Qwen3.8 Omni Flash 908 12,657 8,743 192.840 $0.00608499

Free models and the free router

Five named free endpoints produced pages. Laguna was rate-limited twice. The free router also produced a page and selected Nemotron 3 Ultra for that request. Its result belongs to that selected model and provider at that moment; a later router call can choose something else.

Model Input tokens Output tokens Reasoning within output Generation, s Actual cost, USD
Nemotron 3 Ultra (free) 877 7,108 1,405 298.243 $0.00
Laguna S 2.1 (free) Unavailable Unavailable Unavailable HTTP 429 twice Unavailable
North Mini Code (free) 807 8,335 5,533 120.962 $0.00
Ling 3.0 Flash VL (free) 899 15,785 7,509 210.597 $0.00
Gemma 4 31B (free) 873 3,472 0 97.723 $0.00
Dots3 Note Preview (free) 840 15,545 13,908 166.214 $0.00
OpenRouter Free router 877 5,094 886 219.967 $0.00

Gemma's free request was marked BYOK in the activity log, with zero recorded usage and zero upstream usage. The other successful free requests used OpenRouter's free routing. Free availability and limits can change.

Where does the developer's work begin?

All 21 successful results had one H1 and loaded the existing image. Each contained four FAQ controls, and we verified that an FAQ could be toggled in each result. At 1440 px, none of the pages had horizontal overflow.

At 390 px, eight results stayed within the viewport: Grok, Sonnet, Qwen3 Coder, MiniMax, GLM, North, Ling and Dots. The other 13 measured 402 px wide, a 12 px overflow. That is a small layout correction, but it is still a correction. The screenshots and saved HTML preserve it.

Copy also needs a human check. Qwen3 Coder described ingredients appearing automatically on the shopping list, which the brief had not established. North made a similar claim and turned the instruction about the headline into the literal heading “About deciding dinner before everyone is hungry”. Dots jumped from H2 to H5 in parts of the page. We would correct these before publishing a client page.

These are familiar development decisions: adjust a container that runs past the phone viewport, rewrite a headline and remove a claim the product cannot support. The human part is also where a page gets its character. You may want a quieter hero, a different example meal or a stronger connection between two sections. With the draft already in LiveCanvas, you can make those choices in the visual editor or directly in the HTML.

A follow-up request to the Assistant is another option, with its own generation time and charge. We kept those iterations outside this experiment so the tables describe what the first call delivered. For a client project, the useful measure is how much work that draft saves after you have reviewed it.

Did the AI generate the images too?

No new image was generated in this benchmark. Every successful page used the same existing file, /wp-content/uploads/2026/08/hearthline-meal-planning-2.png. The models returned HTML containing an image reference. We supplied that reference in the prompt.

This distinction matters for both the comparison and the cost. The figures above cover the page-generation calls. They include no separate image-generation request. A model accepting image input does not mean that this endpoint produces image files; the tested catalogue entries listed text output.

If a page needs a new image, that becomes a separate step: create or choose the asset, upload it to WordPress, then give its URL to the Assistant. Its cost and production time should be recorded separately. For this test, reusing the site image kept that part of the brief identical.

Every generated result

Open each comparison below to see the desktop and mobile captures. These images show the top of each first result. Each comparison also has links to the full-page desktop and mobile captures.

Claude Opus 5.5
Claude Opus 5.5 desktop and mobile first result
Claude Opus 5.5. 402 px document width at a 390 px viewport.

Full desktop screenshot   Full mobile screenshot

GPT-6 Astra
GPT-6 Astra desktop and mobile first result
GPT-6 Astra. 402 px document width at a 390 px viewport.

Full desktop screenshot   Full mobile screenshot

Gemini 3.1 Pro Preview
Gemini 3.1 Pro Preview desktop and mobile first result
Gemini 3.1 Pro Preview. 402 px document width at a 390 px viewport.

Full desktop screenshot   Full mobile screenshot

Grok 4.7
Grok 4.7 desktop and mobile first result
Grok 4.7. 390 px without horizontal overflow.

Full desktop screenshot   Full mobile screenshot

Claude Sonnet 4.6
Claude Sonnet 4.6 desktop and mobile first result
Claude Sonnet 4.6. 390 px without horizontal overflow.

Full desktop screenshot   Full mobile screenshot

GPT-6 Sol
GPT-6 Sol desktop and mobile first result
GPT-6 Sol. 402 px document width at a 390 px viewport.

Full desktop screenshot   Full mobile screenshot

MiMo V2.6 Pro
MiMo V2.6 Pro desktop and mobile first result
MiMo V2.6 Pro. 402 px document width at a 390 px viewport.

Full desktop screenshot   Full mobile screenshot

Kimi K2.5
Kimi K2.5 desktop and mobile first result
Kimi K2.5. 402 px document width at a 390 px viewport.

Full desktop screenshot   Full mobile screenshot

Qwen3 Coder
Qwen3 Coder desktop and mobile first result
Qwen3 Coder. 390 px without horizontal overflow.

Full desktop screenshot   Full mobile screenshot

DeepSeek V4.1 Flash
DeepSeek V4.1 Flash desktop and mobile first result
DeepSeek V4.1 Flash. 402 px document width at a 390 px viewport.

Full desktop screenshot   Full mobile screenshot

MiniMax M2.5
MiniMax M2.5 desktop and mobile first result
MiniMax M2.5. 390 px without horizontal overflow.

Full desktop screenshot   Full mobile screenshot

GLM 5.3 Flash
GLM 5.3 Flash desktop and mobile first result
GLM 5.3 Flash. 390 px without horizontal overflow.

Full desktop screenshot   Full mobile screenshot

GPT-6 Luna
GPT-6 Luna desktop and mobile first result
GPT-6 Luna. 402 px document width at a 390 px viewport.

Full desktop screenshot   Full mobile screenshot

Gemini 3.8 Flash
Gemini 3.8 Flash desktop and mobile first result
Gemini 3.8 Flash. 402 px document width at a 390 px viewport.

Full desktop screenshot   Full mobile screenshot

Qwen3.8 Omni Flash
Qwen3.8 Omni Flash desktop and mobile first result
Qwen3.8 Omni Flash. 402 px document width at a 390 px viewport.

Full desktop screenshot   Full mobile screenshot

Nemotron 3 Ultra (free)
Nemotron 3 Ultra (free) desktop and mobile first result
Nemotron 3 Ultra (free). 402 px document width at a 390 px viewport.

Full desktop screenshot   Full mobile screenshot

North Mini Code (free)
North Mini Code (free) desktop and mobile first result
North Mini Code (free). 390 px without horizontal overflow.

Full desktop screenshot   Full mobile screenshot

Ling 3.0 Flash VL (free)
Ling 3.0 Flash VL (free) desktop and mobile first result
Ling 3.0 Flash VL (free). 390 px without horizontal overflow.

Full desktop screenshot   Full mobile screenshot

Gemma 4 31B (free)
Gemma 4 31B (free) desktop and mobile first result
Gemma 4 31B (free). 402 px document width at a 390 px viewport.

Full desktop screenshot   Full mobile screenshot

Dots3 Note Preview (free)
Dots3 Note Preview (free) desktop and mobile first result
Dots3 Note Preview (free). 390 px without horizontal overflow.

Full desktop screenshot   Full mobile screenshot

OpenRouter Free router
OpenRouter Free router desktop and mobile first result
OpenRouter Free router. 402 px document width at a 390 px viewport.

Full desktop screenshot   Full mobile screenshot

The brief behind the comparison

A clear brief gives the model something specific to build. Ours combines the site's offer with the theme colours, the required sections and the rules for editable HTML. We used it unchanged for every call, including the instruction to avoid invented product claims. The few claims that still appeared are a useful reason to review the output.

The complete prompt is included below so you can inspect the experiment or adapt the brief to your own site. Replace the Hearthline details and image path with your own content, keeping the layout requirements that suit your page.

LiveCanvas

AI Assistant

Model
Choose your model
in LiveCanvas
Contexts in this test
Website InformationOffPage TextOff

Paste this into AI Prompts for an empty section in LiveCanvas.

The prompt is unchanged from the benchmark. Before generating on your site, replace the demo image path with a Media Library image of your own. Adapt the Hearthline facts and theme settings for your site.


Rates, capabilities and measurement notes

We checked the OpenRouter catalogue on 23 September 2026. Its listed input/output rates are reference prices; provider routing and caching can change the actual bill. Qwen3 Coder's catalogue entry listed $0.30/$1.00 per million tokens, for example, while the recorded routed call cost $0.00821772. The actual usage charge is the source for the benchmark tables.

The catalogue identified MiniMax M2.5 and Qwen3 Coder as text-input models. The other paid entries listed image input, with additional input types on some models. That is catalogue metadata, not a vision test. We used text prompts throughout.

Catalogue rates and input types for the paid shortlist
Model Input / 1M tokens Output / 1M tokens Listed input types
Claude Opus 5.5 $4 $20 text, image, file
GPT-6 Astra $10 $50 file, image, text
Gemini 3.1 Pro Preview $2 $12 audio, file, image, text, video
Grok 4.7 $1.6 $4.8 text, image, file
Claude Sonnet 4.6 $3 $15 text, image, file
GPT-6 Sol $2 $10 file, image, text
MiMo V2.6 Pro $0.435 $0.87 text, image, video, audio
Kimi K2.5 $0.45 $2.25 text, image
Qwen3 Coder $0.3 $1 text
DeepSeek V4.1 Flash $0.15 $0.6 text, image
MiniMax M2.5 $0.27 $1.08 text
GLM 5.3 Flash $0.15 $0.5 text, image, video
GPT-6 Luna $0.1 $0.5 file, image, text
Gemini 3.8 Flash $0.75 $3.75 text, image, video, file, audio
Qwen3.8 Omni Flash $0.15 $0.47 text, image, audio, video

The available free IDs were nvidia/nemotron-3-ultra-550b-a55b:free and dots-studio/dots-3-note-preview:free; we used those exact catalogue IDs. The model links in the tables identify every requested endpoint.

OpenRouter reports generation time, latency and router latency as separate fields. We preserved those fields separately in the dataset and did not add them into an invented total. We also recorded the time from clicking Go ask AI until completion was observed. Some UI checks had gaps, so those observations are upper bounds rather than precise end-to-end measurements. Saving, review and media preparation are outside the generation figures.

Copy the complete benchmark CSV
category,model,status,attempts,page_id,provider_name,actual_model,cost_usd,input_tokens,output_tokens,reasoning_tokens,cached_tokens,generation_s,latency_s,router_latency_s,ui_last_pending_s,ui_completion_observed_s,desktop_scroll_width,mobile_scroll_width,price_input_per_million,price_output_per_million
Frontier,anthropic/claude-opus-5.5,generated,1,79,Claude Platform on AWS,anthropic/claude-opus-5.5-20260921,0.150148,1322,7243,433,0,56.768,6.223,0.467,,64.363,1440,402,4.0,20.0
Frontier,openai/gpt-6-astra,generated,1,81,OpenAI,openai/gpt-6-astra-20260903,0.22352,822,4306,197,0,73.218,9.515,0.383,73.753,73.89,1440,402,10.0,50.0
Frontier,google/gemini-3.1-pro-preview,generated,1,83,Google,google/gemini-3.1-pro-preview-20260219,0.18209,871,15029,10228,0,102.821,2.993,0.254,103.229,103.38,1440,402,2.0,12.0
Frontier,x-ai/grok-4.7,generated,1,85,xAI,x-ai/grok-4.7-20260916,0.0749456,2084,15207,11747,1152,209.388,0.952,0.268,210.019,210.121,1440,390,1.6,4.8
Default candidates,anthropic/claude-sonnet-4.6,generated,1,87,Claude Platform on AWS,anthropic/claude-4.6-sonnet-20260217,0.116214,913,7565,0,0,93.129,0.941,0.254,93.563,93.689,1440,390,3.0,15.0
Default candidates,openai/gpt-6-sol,generated,1,89,OpenAI,openai/gpt-6-sol-20260922,0.038534,822,3689,825,0,38.428,4.721,0.18,41.239,41.813,1440,402,2.0,10.0
Open weight,xiaomi/mimo-v2.6-pro,generated,1,91,Xiaomi,xiaomi/mimo-v2.6-pro-20260921,0.00459012,844,4854,546,0,91.861,5.911,1.068,93.967,94.093,1440,402,0.435,0.87
Open weight,moonshotai/kimi-k2.5,generated,1,93,SiliconFlow,moonshotai/kimi-k2.5-0127,0.0135846,828,5872,1003,0,195.962,1.264,0.166,196.107,220.794,1440,402,0.45,2.25
Open weight,qwen/qwen3-coder,generated,1,95,Google,qwen/qwen3-coder-480b-a35b-07-25,0.00821772,846,4462,0,5,22.334,1.327,0.178,23.552,23.649,1440,390,0.3,1.0
Open weight,deepseek/deepseek-v4.1-flash,generated,1,97,DeepInfra,deepseek/deepseek-v4.1-flash-20260910,0.0041923,860,9695,4660,0,180.958,0.684,0.164,181.986,182.107,1440,402,0.15,0.6
Open weight,minimax/minimax-m2.5,generated,1,99,AtlasCloud,minimax/minimax-m2.5-20260211,0.021603925,835,17798,13970,0,363.144,0.889,0.154,318.836,390.688,1440,390,0.27,1.08
Open weight,z-ai/glm-5.3-flash,generated,1,101,Together,z-ai/glm-5.3-flash-20260826,0.01026135,829,20274,15759,0,240.867,0.384,0.172,241.832,241.938,1440,390,0.15,0.5
Maximum savings,openai/gpt-6-luna,generated,1,103,OpenAI,openai/gpt-6-luna-20260922,0.0022517,822,4339,873,0,47.557,17.031,0.177,45.659,66.324,1440,402,0.1,0.5
Maximum savings,google/gemini-3.8-flash,generated,1,105,Google,google/gemini-3.8-flash-20260902,0.032172,871,8405,2362,0,82.938,1.923,0.279,84.075,84.172,1440,402,0.75,3.75
Budget multimodal,qwen/qwen3.8-omni-flash,generated,1,107,Alibaba,qwen/qwen3.8-omni-flash-20260918,0.00608499,908,12657,8743,0,192.84,0.643,0.171,194.252,194.348,1440,402,0.15,0.47
Free,nvidia/nemotron-3-ultra-550b-a55b:free,generated,1,109,Nvidia,nvidia/nemotron-3-ultra-550b-a55b-20260604:free,0,877,7108,1405,0,298.243,5.33,0.172,295.703,319.874,1440,402,0.0,0.0
Free,poolside/laguna-s-2.1:free,error_429,2,111,,,,,,,,,,,0.0,8.468,,,0.0,0.0
Free,cohere/north-mini-code:free,generated,1,113,Cohere,cohere/north-mini-code-20260617:free,0,807,8335,5533,0,120.962,0.333,0.179,122.761,123.427,1440,390,0.0,0.0
Free,inclusionai/ling-3.0-flash-vl:free,generated,1,115,Novita,inclusionai/ling-3.0-flash-vl-20260910:free,0,899,15785,7509,3,210.597,1.072,0.172,208.368,229.717,1440,390,0.0,0.0
Free,google/gemma-4-31b-it:free,generated,1,117,Google AI Studio,google/gemma-4-31b-it-20260402:free,0,873,3472,0,0,97.723,1.813,0.398,96.015,118.949,1440,402,0.0,0.0
Free,dots-studio/dots-3-note-preview:free,generated,1,119,AtlasCloud,dots-studio/dots-3-note-preview-20260813:free,0,840,15545,13908,0,166.214,1.484,0.16,48.775,219.476,1440,390,0.0,0.0
Free,openrouter/free,generated,1,121,Nvidia,nvidia/nemotron-3-ultra-550b-a55b-20260604:free,0,877,5094,886,0,219.967,0.725,0.181,205.174,231.009,1440,402,0.0,0.0

The results are first attempts from one staging site, on one date, with one additional attempt for the failed Laguna endpoint. The saved dataset includes the actual provider and canonical model name for each successful call.

Provider and model-family logos: Lobe Icons. Brand names identify the models tested.