InsightsPublished 5 min read

GLM-5.3-Flash: an AI that can check what is on screen

The update adds visual input, so software can use screenshots and documents while working. We explain the possible uses, reported test results and costs.

By Monolith

An architecture diagram on paper with a highlighted decision point
A site map on paper with one screen circled for a closer check.

Z.ai released GLM-5.3-Flash on August 26, 2026. The main addition is the ability to work with visual material. Connected software can show it a web page, ask it to change the code and show it the result again. That creates a useful possibility for checking layouts, although it still needs testing on the work you actually do.1

Show the problem instead of describing it

Imagine a mobile webpage where the headline covers the booking button. A written description might leave out the screen size or the exact overlap. A screenshot shows what the customer sees. With access to a copy of the website, connected software could ask the model to propose a fix, apply it and inspect a fresh screenshot.

That is the practical promise of visual input: the AI can use appearance as part of the task. It could help with a layout review or reading a chart. The result still needs an ordinary check, such as clicking the booking button, because a page can look fixed while its behavior remains wrong.

What changed versus GLM-5.2

Earlier GLM-5.2 work was text-based. Flash can also receive images, video and files, so you can show it a problem instead of describing every detail. It still replies with text. If you want it to produce an editable slide deck or repair a website, another part of the software must carry out those instructions.1

The model’s internal size is less useful to most buyers than the work it can complete. Parameters are learned settings inside the model; counting them is not a measure of how well it understands your brand or checks a spreadsheet. Keep the design details as background, then judge a trial on its result.

DetailWhat it means
Design320 billion parameters in total, with 18 billion active at a time.
Working spaceOne-million-token context window.
Generated responseUp to 128K output tokens.
Technical reference for the model described by Z.ai.14

How it compares with GLM-5.2

MeasureGLM-5.2GLM-5.3-Flash
DeepSWE v1.146.263.4
AutomationBench26.248.8
Z.ai-published comparisons in the Flash documentation. These are not independent Monolith measurements.1

The two rows test different jobs. DeepSWE asks the system to repair software, while AutomationBench tests work involving several steps. Z.ai reports improvement over GLM-5.2 on both. This makes Flash a candidate for a trial, but it does not tell you how often it will succeed on your own files.

DetailWhat it means
DeepSWE setupmini-swe-agent supporting software, a six-hour time limit and 400K context.
AutomationBench version1.0.6. Scores from another version or setup are not a like-for-like comparison.
Comparison conditions to retain when reading the scores.4

Practical uses worth evaluating

Try one visible website problem first, such as a heading cut off on a phone. Supply a screenshot and a copy of the site, then ask for the change and a new screenshot. Check whether the page also works with a keyboard and whether its buttons still do the right thing. Looking correct is only part of working correctly.

Another trial is a presentation or spreadsheet built from approved source material. Ask for editable files, then check both the appearance and the facts. For a spreadsheet, verify formulas as well as the displayed totals. These are suggested trials based on the documented capabilities, not reports of Monolith testing this model.1

For a presentation trial, supply the approved facts and ask for one slide that explains them. Check whether the message is easy to follow before asking for a whole deck. You might discover that the model needs a clearer audience description or an example of your preferred layout. A small first deliverable makes those corrections easier.

Availability, limits and pricing

The published usage rates are lower than the GLM-5.2 rates shown here. That could reduce the AI processing part of a job, if the model produces results you can use. Include repeat attempts and human corrections when comparing total cost.

DetailWhat it means
GLM-5.3-FlashInput $0.15; eligible cached input $0.03; output $0.50.
GLM-5.2 comparisonInput $1.40; output $4.40.
Model name for softwareglm-5.3-flash.
Published USD API rates per million tokens.2

Flash always uses its thinking mode; the documentation does not offer a switch to turn it off. It is also offered through the GLM Coding Plan, whose usage allowance is separate from pay-as-you-go billing. The downloadable model is published under the MIT license, but running it yourself still requires suitable equipment and technical support. This article does not include a hardware test.13

Start with a workflow whose result is easy to check. A useful trial might repair one page, extract a small document set, or produce a deck from approved source material. Record accepted output and review time alongside the API bill. That gives a more useful answer than replacing the default model because its launch chart looks better.

Terms used in this article

TermPlain-language meaning
TokenA small unit of information counted by the model. Usage fees often depend on how many are read and generated.
Input / outputThe information you send / the response the model generates.
Cached inputEligible information stored for reuse, sometimes charged at a lower rate.
BenchmarkA defined test. Its score depends on the tasks, settings and supporting tools.
Context windowHow much information fits into one request; it is not a guarantee of perfect recall.
A quick guide to terms used in this article.
Questions, answered
What changed from GLM-5.2?
GLM-5.3-Flash adds native visual input while retaining a million-token context scale. It generates text and uses tools for other deliverables.1
How much does the API cost?
Z.ai lists $0.15 input, $0.03 cached input and $0.50 output per million tokens. That excludes other workflow costs.2
Can GLM-5.3-Flash be self-hosted?
The published checkpoint carries an MIT license. Infrastructure requirements still need to be assessed for the intended workload.34
Sources

Read for this feature. The numbers match the markers in the text.

  1. Z.ai GLM-5.3-Flash documentationdocs.z.ai
  2. Z.ai API pricingdocs.z.ai
  3. GLM-5.3-Flash MIT licensehuggingface.co
  4. GLM-5.3-Flash model cardhuggingface.co
Where this leads

Want to apply this to a project? Here is the related service.

Free download

25 AI instructions to try on everyday work

Prompts are instructions you give an AI tool. This PDF includes 25 examples for inquiries, marketing and admin. Adapt them to your task and check the results.

Enter your email to access the PDF. This form does not sign you up for a marketing sequence.